CronJobs

devops-sre jobs

Principal SRE - AI Inference

Cerebras · Sunnyvale, CA

hybridprincipalPosted Jul 6, 2026SREobservabilitySLO/SLIincident responseorchestrationcapacity managementautomationAI inference

Apply on the employer site

About this role

## Principal SRE — AI Inference **Cerebras Systems** builds the world’s largest AI chip (Wafer-Scale Engine), enabling industry-leading training and inference speeds—over 10× faster than GPU-based hyperscale cloud inference services. ### About the Role We’re building a high-performance **SRE** function to support one of the world’s fastest-growing **AI inference services**. As a **Principal SRE**, you’ll define and drive the technical architecture for scaling our inference fleet through: - **Self-service delivery** - **Shared observability** - **Capacity orchestration** - **Rollout safety** - **Operational automation** This role begins with **2–3 weeks of hands-on operational immersion** to build deep context on the current stack, production pain points, and high-stakes workflows. You’ll then shift to architecting the “tomorrow” layer: a unified **capacity management and production control plane** for reliable capacity planning, workload placement, rollout safety, validation, and operational decision-making across large-scale inference infrastructure. **No 24/7 on-call rotations.** ### Key Responsibilities - Define and implement a robust strategy for delivering and running software reliably and at scale across multiple datacenters and cloud-based solutions. - Architect **self-service platforms** and internal tooling so product teams, external customers, and cluster operators can safely trigger and observe critical workflows with minimal handoffs. - Define and evolve reliability practices for inference workloads, including: - **SLOs/SLIs** (latency, throughput, accuracy stability) - **Error budgets** - **Blameless postmortems** - **Chaos testing** - **Capacity forecasting** across multi-datacenter and on-prem environments - Mentor senior SREs, support critical incident escalations, and prioritize high-leverage automation based on production pain points. - Measure and drive impact using clear metrics (toil reduction, deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows). ### Required Experience & Skills - **15+ years** in SRE, infrastructure engineering, or platform engineering; proven record of setting technical direction and delivering reliability improvements at large scale (FAANG/hyperscaler/frontier AI or similarly demanding environments). - Deep experience with large-scale compute fleets, internal control planes, schedulers, orchestration systems, capacity management, and reliability automation. - Experience defining and driving cross-team architecture for production control planes, capacity orchestration, fleet management, or self-service infrastructure platforms with clear operational ownership. - Strong judgment to converge fragmented workflows, tools, and teams into coherent architectures that improve reliability and operational leverage. - Ability to lead complex, ambiguous technical programs end to end; influence senior cross-functional stakeholders; mentor senior engineers; and communicate technical strategy clearly. - Hands-on experience with production observability, incident response, and **SLO-based reliability management** across metrics, logs, traces, alerting, dashboards, and operational review loops. ### Nice-to-Haves - Experience with **Bazel** or other large-scale build systems in production. - Background in **AI/ML inference systems** (model serving runtimes, disaggregated inference, GPU orchestration, latency/accuracy SLOs, drift monitoring). - Experience with predicti

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord