Principal SRE - AI Inference
Cerebras · Sunnyvale, CA
About this role
## Principal SRE — AI Inference **Cerebras Systems** builds the world’s largest AI chip (Wafer-Scale Engine), enabling industry-leading training and inference speeds—over 10× faster than GPU-based hyperscale cloud inference services. ### About the Role We’re building a high-performance **SRE** function to support one of the world’s fastest-growing **AI inference services**. As a **Principal SRE**, you’ll define and drive the technical architecture for scaling our inference fleet through: - **Self-service delivery** - **Shared observability** - **Capacity orchestration** - **Rollout safety** - **Operational automation** This role begins with **2–3 weeks of hands-on operational immersion** to build deep context on the current stack, production pain points, and high-stakes workflows. You’ll then shift to architecting the “tomorrow” layer: a unified **capacity management and production control plane** for reliable capacity planning, workload placement, rollout safety, validation, and operational decision-making across large-scale inference infrastructure. **No 24/7 on-call rotations.** ### Key Responsibilities - Define and implement a robust strategy for delivering and running software reliably and at scale across multiple datacenters and cloud-based solutions. - Architect **self-service platforms** and internal tooling so product teams, external customers, and cluster operators can safely trigger and observe critical workflows with minimal handoffs. - Define and evolve reliability practices for inference workloads, including: - **SLOs/SLIs** (latency, throughput, accuracy stability) - **Error budgets** - **Blameless postmortems** - **Chaos testing** - **Capacity forecasting** across multi-datacenter and on-prem environments - Mentor senior SREs, support critical incident escalations, and prioritize high-leverage automation based on production pain points. - Measure and drive impact using clear metrics (toil reduction, deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows). ### Required Experience & Skills - **15+ years** in SRE, infrastructure engineering, or platform engineering; proven record of setting technical direction and delivering reliability improvements at large scale (FAANG/hyperscaler/frontier AI or similarly demanding environments). - Deep experience with large-scale compute fleets, internal control planes, schedulers, orchestration systems, capacity management, and reliability automation. - Experience defining and driving cross-team architecture for production control planes, capacity orchestration, fleet management, or self-service infrastructure platforms with clear operational ownership. - Strong judgment to converge fragmented workflows, tools, and teams into coherent architectures that improve reliability and operational leverage. - Ability to lead complex, ambiguous technical programs end to end; influence senior cross-functional stakeholders; mentor senior engineers; and communicate technical strategy clearly. - Hands-on experience with production observability, incident response, and **SLO-based reliability management** across metrics, logs, traces, alerting, dashboards, and operational review loops. ### Nice-to-Haves - Experience with **Bazel** or other large-scale build systems in production. - Background in **AI/ML inference systems** (model serving runtimes, disaggregated inference, GPU orchestration, latency/accuracy SLOs, drift monitoring). - Experience with predicti
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.