Sr. Staff Software Engineer, Inference Platform
Cerebras · Sunnyvale, CA
About this role
## Sr. Staff Software Engineer, Inference Platform ### About Cerebras Cerebras builds the world’s largest AI chip—an architecture designed to deliver industry-leading training and inference speeds (over 10× faster than GPU-based hyperscale cloud inference services). Cerebras works with leading model labs, global enterprises, and AI-native startups. ### About the Role Cerebras is hiring a **Staff Engineer** to lead and contribute to projects on the **Inference Platform** team. This team primarily owns the orchestration layer that runs inference on Cerebras datacenter clusters—connecting cloud components with machine learning services. You’ll be on the front line of problems that haven’t been solved yet, driving solutions across **Kubernetes operators**, **service security policies**, and **CI/CD**. ### Responsibilities - **Design, develop, test, and maintain** production software across testing, continuous development, observability, security, networking, debugging, and productionization. - **Increase the effectiveness of senior engineers** through design feedback, pairing, and clear technical standards. - **Platform direction:** help shape technical direction for the Inference Platform, including Kubernetes custom resource definitions, failure domains, service boundaries, and system evolution; own roadmaps for major technical areas. - **Reliability & performance:** architect active-active systems with rapid failover, graceful degradation, and clear SLOs; improve latency, throughput, capacity efficiency, and resilience under unpredictable demand. - **Execution on critical paths:** write and review production code in the most important parts of the platform; make high-consequence architectural decisions; set the technical bar through design and code reviews. - **Production leadership:** lead hardest production issues and cross-system bottlenecks; drive observability, incident response, capacity planning, and post-incident improvements with strong operational rigor. - **Technical influence:** partner with ML, Product, Infrastructure, and Cloud teams to translate requirements into scalable system designs and drive alignment on shared technical decisions. ### Skills & Qualifications - **8+ years** of software engineering experience, including substantial individual contributor experience building and operating large-scale distributed systems or cloud infrastructure. - Deep expertise in **distributed systems architecture**, ideally with **Kubernetes**. - Strong track record making sound architectural decisions for highly available, latency-sensitive systems at scale. - Experience with **security** (certificates, TLS, mTLS). - Experience optimizing **latency, throughput, and efficiency** in high-QPS systems (TTFT and tail-latency reduction is a plus). - Proficiency in backend/systems languages such as **Go or C++**, with the expectation you can contribute production code directly. - Experience designing **observability and reliability practices** (metrics, logging, tracing, alerting, incident response, SLO-driven operations). - Ability to influence senior engineers and cross-functional partners through technical credibility, communication, and judgment. ### Preferred Skills & Qualifications - Experience with **ML inference infrastructure**, model serving systems, or **GPU-accelerated workloads**. ### Location **Sunnyvale or Toronto preferred** ### Why Join Cerebras People who are serious about software make their own hardware. Cerebras
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.