Senior Staff AI Accelerator Performance Architect
Cerebras · Sunnyvale, CA
About this role
## Senior Staff AI Accelerator Performance Architect **Cerebras Systems** builds the world’s largest AI chip—**56× larger than GPUs**—enabling **industry-leading training and inference speeds** (over **10× faster** than GPU-based hyperscale cloud inference). Cerebras works with leading model labs, global enterprises, and AI-native startups. ### What you’ll do - Own and evolve **performance models** and **modeling methodologies** for next-generation accelerator and system architectures. - Build and extend **analytical, simulation-based, or trace-driven models** across workloads, architectural features, and product generations. - Analyze AI workloads from **individual kernels** to **end-to-end inference and training** to determine where **time, bandwidth, compute, and capacity** are spent. - Identify **hardware and software bottlenecks** and quantify opportunities to improve **latency, throughput, utilization, and energy efficiency**. - Evaluate proposed architectural features and estimate expected **performance return** across representative workloads. - Study how models and kernels map onto underlying **compute, memory, and communication** architecture. - Partner with **architecture, compiler, kernel, runtime, and systems** teams to evaluate alternative mappings and optimizations. - Develop workload projections and **competitive performance analyses** with transparent assumptions. - Translate complex performance results into clear **architectural and product recommendations**. - Improve modeling methodology, validation, and correlation with **RTL, emulation, and silicon measurements**. - Help define representative workloads, performance targets, and success criteria for future products. ### What we’re looking for - **7+ years** experience in performance analysis, performance modeling, or architecture exploration for **CPUs, GPUs, AI accelerators**, or other high-performance computing systems. - Strong understanding of hardware architecture developed through hardware, compiler, kernel, runtime, or system-performance work. - Experience building analytical/simulation/trace-driven performance models using **Python, C++**, or similar. - Solid understanding of processor architecture, **memory systems**, interconnects, parallel execution, and hardware resource constraints. - Ability to connect **kernel-level behavior** to **end-to-end application/system** performance. - Experience profiling workloads, forming performance hypotheses, and validating with quantitative evidence. - Understanding of how software mapping and programmability affect realized hardware performance. - Ability to clearly communicate assumptions, uncertainty, bottlenecks, and recommendations. - **MS or PhD** in EE, CE, CS, or equivalent practical experience. ### Particularly relevant experience - Performance analysis of **transformer inference/training** workloads. - Experience with **attention**, **GEMM/GEMV**, collective communication, **Mixture-of-Experts**, quantization, or memory-capacity-constrained execution. - Kernel optimization, compiler performance, runtime scheduling, or distributed accelerator systems. - Model validation using **RTL simulation, emulation, FPGA prototypes, or silicon measurements**. - Competitive analysis of AI accelerators and large-scale AI systems. ### Role focus This is a **performance and architecture** role—not a production RTL-design position. You’ll reason about microarchitecture and work with architecture/RTL/physical-design teams,
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.