Staff Software Engineer, Inference Cloud
Cerebras · Sunnyvale, CA
About this role
## Staff Software Engineer, Inference Cloud ### About the Role Cerebras builds AI hardware and the cloud infrastructure to deliver **industry-leading training and inference speeds**—enabling **real-time iteration** and more agentic computation. We’re hiring a **Staff Engineer** to own major areas of the **Inference Cloud Platform**, focusing on the cloud layer behind our Inference Service. This is a hands-on individual contributor role tackling hard distributed systems problems such as **multi-region traffic architecture**, **graceful degradation under bursty workloads**, **high-QPS performance**, and the **operating model** for a platform that must stay fast and available under load. ### Responsibilities - **Platform Direction:** Shape technical direction for the Inference Cloud Platform (multi-region topology, failure domains, service boundaries, and system evolution) and own roadmaps for major areas. - **Core Cloud Systems:** Design and build critical components including **service discovery, request routing, load balancing, caching, batching, and traffic management** for inference workloads. - **Reliability & Performance:** Architect **active-active** systems with rapid failover, graceful degradation, and clear **SLOs**; improve latency, throughput, capacity efficiency, and resilience. - **Traffic Control & Service Tiers:** Define mechanisms for **admission control, quota management, rate limiting**, and differentiated **quality of service**. - **Execution on Critical Paths:** Write and review production code in the most important parts of the platform; make high-consequence architectural decisions and set technical standards. - **Production Leadership:** Lead hardest production issues and cross-system bottlenecks; drive observability, incident response, capacity planning, and post-incident improvements. - **Technical Influence:** Partner with ML, Product, Infrastructure, and Platform teams to translate requirements into scalable system designs and align on shared decisions. - **Mentorship:** Improve senior engineer effectiveness through design feedback, pairing, and clear technical standards. ### Skills & Qualifications - **8+ years** of software engineering experience, with substantial individual contributor experience building and operating large-scale distributed systems or cloud infrastructure. - Deep expertise in **distributed systems architecture** in cloud environments (networking, compute orchestration, container platforms, multi-region production services). - Proven ability to make sound architectural decisions for **highly available, latency-sensitive** systems at scale. - Experience optimizing **latency, throughput, and efficiency** in **high-QPS** systems (experience with **TTFT** and **tail-latency reduction** is a plus). - Strong proficiency in backend/systems languages such as **Go, C++, or Python**; ability to contribute production code directly. - Experience designing **observability and reliability** practices (metrics, logging, tracing, alerting, incident response, SLO-driven operations). - Ability to influence senior engineers and cross-functional partners through technical credibility, communication, and judgment. ### Preferred Skills - Experience with **ML inference infrastructure**, model serving systems, or **GPU-accelerated workloads**. ### Why Join Cerebras People who are serious about software make their own hardware. With rapid growth and many model releases, you’ll help build a breakthrough AI pla
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.