CronJobs

devops-sre jobs

Senior Staff AI Accelerator Performance Architect

Cerebras · Sunnyvale, CA

onsitesenior$175,000–$275,000Posted Sep 17, 2026PythonC++transformersGEMMGEMVquantizationRTL simulationFPGA

Apply on the employer site

About this role

## Senior Staff AI Accelerator Performance Architect **Cerebras Systems** builds the world’s largest AI chip—**56× larger than GPUs**—enabling **industry-leading training and inference speeds** (over **10× faster** than GPU-based hyperscale cloud inference). Cerebras works with leading model labs, global enterprises, and AI-native startups. ### What you’ll do - Own and evolve **performance models** and **modeling methodologies** for next-generation accelerator and system architectures. - Build and extend **analytical, simulation-based, or trace-driven models** across workloads, architectural features, and product generations. - Analyze AI workloads from **individual kernels** to **end-to-end inference and training** to determine where **time, bandwidth, compute, and capacity** are spent. - Identify **hardware and software bottlenecks** and quantify opportunities to improve **latency, throughput, utilization, and energy efficiency**. - Evaluate proposed architectural features and estimate expected **performance return** across representative workloads. - Study how models and kernels map onto underlying **compute, memory, and communication** architecture. - Partner with **architecture, compiler, kernel, runtime, and systems** teams to evaluate alternative mappings and optimizations. - Develop workload projections and **competitive performance analyses** with transparent assumptions. - Translate complex performance results into clear **architectural and product recommendations**. - Improve modeling methodology, validation, and correlation with **RTL, emulation, and silicon measurements**. - Help define representative workloads, performance targets, and success criteria for future products. ### What we’re looking for - **7+ years** experience in performance analysis, performance modeling, or architecture exploration for **CPUs, GPUs, AI accelerators**, or other high-performance computing systems. - Strong understanding of hardware architecture developed through hardware, compiler, kernel, runtime, or system-performance work. - Experience building analytical/simulation/trace-driven performance models using **Python, C++**, or similar. - Solid understanding of processor architecture, **memory systems**, interconnects, parallel execution, and hardware resource constraints. - Ability to connect **kernel-level behavior** to **end-to-end application/system** performance. - Experience profiling workloads, forming performance hypotheses, and validating with quantitative evidence. - Understanding of how software mapping and programmability affect realized hardware performance. - Ability to clearly communicate assumptions, uncertainty, bottlenecks, and recommendations. - **MS or PhD** in EE, CE, CS, or equivalent practical experience. ### Particularly relevant experience - Performance analysis of **transformer inference/training** workloads. - Experience with **attention**, **GEMM/GEMV**, collective communication, **Mixture-of-Experts**, quantization, or memory-capacity-constrained execution. - Kernel optimization, compiler performance, runtime scheduling, or distributed accelerator systems. - Model validation using **RTL simulation, emulation, FPGA prototypes, or silicon measurements**. - Competitive analysis of AI accelerators and large-scale AI systems. ### Role focus This is a **performance and architecture** role—not a production RTL-design position. You’ll reason about microarchitecture and work with architecture/RTL/physical-design teams,

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord