CronJobs

backend jobs

Director/Sr. Manager, AI Inference Model Scaling

Cerebras · Sunnyvale, CA

hybridseniorPosted Jul 23, 2026PythonC++LLVMMLIRXLATVMTorch FXPyTorch

Apply on the employer site

About this role

## About Cerebras Cerebras Systems builds the world’s largest AI chip—56x larger than GPUs—enabling industry-leading training and inference speeds (10x faster than GPU-based hyperscale cloud inference services). This performance leap unlocks real-time iteration and more agentic computation for smarter AI applications. ## About the Team The **Inference Model Scaling** team enables state-of-the-art foundation models and generative AI workloads to run efficiently on Cerebras’ **Wafer-Scale Engine (WSE)**. You’ll help build: - Compiler frontend - Model transformation pipeline - Graph optimization infrastructure - High-performance kernel enablement - Runtime integration The team works across ML frameworks, compiler technologies, distributed systems, hardware architecture, and model optimization—partnering with hardware architects, runtime engineers, cloud platform teams, AI researchers, and strategic customers. ## About the Role Cerebras is seeking an experienced engineering leader to **build and scale the Inference Model Scaling organization**. You will define the technical vision, organizational strategy, and execution roadmap for a globally distributed engineering team enabling the latest foundation models on Cerebras hardware. You will lead engineering for **ML model compilation and optimization** as well as **development of high-performance kernels**, partnering across compiler, runtime, cloud infrastructure, hardware architecture, product management, and AI research. ## Responsibilities ### Technical Leadership - Define the technical roadmap and strategy - Establish technical direction across multiple teams and engineering leaders - Lead design reviews and establish engineering standards - Drive support for emerging LLM architectures and inference workloads ### Team Leadership - Hire, mentor, and grow a high-performing engineering team - Develop future technical leaders and managers - Drive organizational planning, headcount strategy, and investment priorities - Foster an engineering culture focused on execution, quality, and innovation - Scale engineering processes while maintaining execution velocity ### Cross-Functional Collaboration - Partner with Cloud Platform, ML, and Hardware teams for end-to-end service enablement (Cloud and On-Premise) - Work with Product Management to prioritize model enablement and customer needs - Collaborate with customers and solution architects on new model bring-up - Influence future hardware/software co-design through ML model enablement and optimization insights ### Delivery & Execution - Own planning, prioritization, and execution across multiple concurrent initiatives - Balance rapid model support with long-term ML Compiler architecture - Drive predictable delivery for strategic customer commitments ## Required Qualifications - BS, MS, or PhD in Computer Science, Computer Engineering, or related field - 12+ years building compiler, ML systems, or infrastructure software - 5+ years leading engineering teams - Deep experience with modern compiler infrastructure (LLVM, MLIR, XLA, TVM, Torch FX, or similar) - Strong understanding of graph compilation and optimization - Experience with Python and C++ - Experience delivering production-quality software - Strong communication and cross-functional leadership skills ## Preferred Qualifications - Experience building compiler frontends for AI accelerators - Experience supporting PyTorch, JAX, TensorFlow, or ONNX - Experience with LLM inference or tr

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord