CronJobs

backend jobs

Staff Software Engineer, Inference Cloud

Cerebras · Sunnyvale, CA

onsiteseniorPosted Jul 12, 2024PythonGoC++distributed-systemscloud-infrastructureobservabilitySLOsload-balancing

Apply on the employer site

About this role

## Staff Software Engineer, Inference Cloud ### About the Role Cerebras builds AI hardware and the cloud infrastructure to deliver **industry-leading training and inference speeds**—enabling **real-time iteration** and more agentic computation. We’re hiring a **Staff Engineer** to own major areas of the **Inference Cloud Platform**, focusing on the cloud layer behind our Inference Service. This is a hands-on individual contributor role tackling hard distributed systems problems such as **multi-region traffic architecture**, **graceful degradation under bursty workloads**, **high-QPS performance**, and the **operating model** for a platform that must stay fast and available under load. ### Responsibilities - **Platform Direction:** Shape technical direction for the Inference Cloud Platform (multi-region topology, failure domains, service boundaries, and system evolution) and own roadmaps for major areas. - **Core Cloud Systems:** Design and build critical components including **service discovery, request routing, load balancing, caching, batching, and traffic management** for inference workloads. - **Reliability & Performance:** Architect **active-active** systems with rapid failover, graceful degradation, and clear **SLOs**; improve latency, throughput, capacity efficiency, and resilience. - **Traffic Control & Service Tiers:** Define mechanisms for **admission control, quota management, rate limiting**, and differentiated **quality of service**. - **Execution on Critical Paths:** Write and review production code in the most important parts of the platform; make high-consequence architectural decisions and set technical standards. - **Production Leadership:** Lead hardest production issues and cross-system bottlenecks; drive observability, incident response, capacity planning, and post-incident improvements. - **Technical Influence:** Partner with ML, Product, Infrastructure, and Platform teams to translate requirements into scalable system designs and align on shared decisions. - **Mentorship:** Improve senior engineer effectiveness through design feedback, pairing, and clear technical standards. ### Skills & Qualifications - **8+ years** of software engineering experience, with substantial individual contributor experience building and operating large-scale distributed systems or cloud infrastructure. - Deep expertise in **distributed systems architecture** in cloud environments (networking, compute orchestration, container platforms, multi-region production services). - Proven ability to make sound architectural decisions for **highly available, latency-sensitive** systems at scale. - Experience optimizing **latency, throughput, and efficiency** in **high-QPS** systems (experience with **TTFT** and **tail-latency reduction** is a plus). - Strong proficiency in backend/systems languages such as **Go, C++, or Python**; ability to contribute production code directly. - Experience designing **observability and reliability** practices (metrics, logging, tracing, alerting, incident response, SLO-driven operations). - Ability to influence senior engineers and cross-functional partners through technical credibility, communication, and judgment. ### Preferred Skills - Experience with **ML inference infrastructure**, model serving systems, or **GPU-accelerated workloads**. ### Why Join Cerebras People who are serious about software make their own hardware. With rapid growth and many model releases, you’ll help build a breakthrough AI pla

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord