CronJobs

devops-sre jobs

Staff Site Reliability Engineer – Automation and Platform

Cerebras · Sunnyvale, CA

hybridstaffPosted Oct 3, 2025KubernetesGitOpsArgoCDPrometheusLokiTempoMimirBazel

Apply on the employer site

About this role

## About the Role Cerebras Systems is building a high-performance SRE function to support one of the world’s fastest-growing AI inference services, powered by the Wafer-Scale Engine (WSE). As a **Staff Site Reliability Engineer (Automation and Platform)**, you’ll lead efforts to eliminate toil at scale by building **self-service delivery pipelines** and **shared observability tooling**. This role begins with ~**1 month of hands-on operational immersion** to learn the current stack, production pain points, and high-stakes workflows. After that, your primary focus shifts to architecting and delivering the “tomorrow” layer: **declarative, GitOps-driven CD for model releases**, plus **capacity provisioning** and **cluster upgrades**. Success in your first year means enabling core teams, product managers, external customers, and cluster stakeholders to operate with **fully self-service workflows** and **strong reliability guarantees**. You’ll partner with an early-career SRE sub-team to automate their toil and mentor them as platform engineers. **On-call:** This role does **not** require 24/7 on-call rotations. --- ## Key Responsibilities - Define and implement a robust strategy for delivering and running software reliably and at scale across multiple datacenters and cloud-based solutions. - Architect self-service platforms and internal tooling so product teams, external customers, and cluster operators can safely trigger and observe critical workflows with minimal handoffs. - Define and evolve reliability practices for inference workloads, including: - **SLOs/SLIs** for latency, throughput, and accuracy stability - **Error budgets** - **Blameless postmortems** - **Chaos testing** - **Capacity forecasting** across multi-datacenter and on-prem environments - Mentor mid-level SREs, support critical incident escalations, and prioritize high-leverage automation based on production pain points. - Drive measurable impact using metrics such as **toil reduction**, **deployment velocity**, **SLO compliance**, **MTTR**, and adoption of self-service workflows. --- ## Required Experience & Skills - **8+ years** in SRE, infrastructure engineering, or platform engineering, with a strong record of improving automation and reliability at large scale in FAANG, hyperscaler, or similarly demanding environments. - Deep expertise operating **large-scale heterogeneous clusters** with a proprietary cloud control plane. - Proven track record designing and delivering **CI/CD or GitOps** systems using **Argo CD** or similar tools, with strong safety and observability. - Hands-on experience with observability systems such as **Loki, Tempo, Mimir, and Prometheus**. - Ability to lead complex end-to-end projects, influence cross-functional stakeholders, and communicate technical direction clearly. --- ## Nice-to-Haves - Experience with **Bazel** or other large-scale build systems in production. - Background in **AI/ML inference systems**, including model serving runtimes, GPU/wafer-scale orchestration, latency & accuracy SLOs, or drift monitoring. - Prior work on **predictive autoscaling**, **chaos engineering**, or **cost-aware capacity planning** for compute-intensive workloads. --- ## Location - **SF Bay Area** - **Toronto** --- ## Why Join Cerebras People who are serious about software make their own hardware. Cerebras is building a breakthrough architecture that’s unlocking new opportunities for the AI industry—backed by rapid growth and freq

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord