CronJobs

devops-sre jobs

Site Reliability Engineering Lead

Graphcore · Austin, Texas, United States

remoteseniorPosted Sep 29, 2026LinuxKubernetesSLOsObservabilityAutomationDistributed systemsNetworkingStorage

Apply on the employer site

About this role

**About Graphcore** Graphcore builds made-for-AI compute hardware and software, helping AI researchers develop advanced models and enabling companies worldwide to put AI at the heart of their business. Graphcore recently joined SoftBank Group, bringing large and ongoing investment from a leading backer of innovative AI companies. **Job Summary** Graphcore is seeking an experienced **Site Reliability Engineering (SRE) leader** to build and lead a new SRE organization responsible for the **production operations** of a rapidly scaling **AI supercomputing platform**. This environment combines highly customized compute, high-performance networking, storage, and supporting infrastructure, and will grow through multiple deployment phases. This is a rare opportunity to establish the reliability function for a new platform from the ground up—taking it through **production launch, stabilization, and scale**. The platform and its operational model are being developed in parallel, with a goal of supporting a **24x7x365** production service with stringent availability requirements. This is **not purely managerial**. During development and early production phases, the SRE Manager will work directly with engineering teams, build deep technical understanding of the platform, and participate in troubleshooting and incident response. Over time, success will increasingly mean building the people, processes, automation, tooling, and operational discipline so the organization can operate effectively without relying on you for day-to-day escalation. **Responsibilities and Duties** - **Build the SRE Organization** - Establish the team from initial formation through full **24x7x365** production operations (roles, interviewing/hiring, career expectations, and developing future technical leaders). - Mentor engineers and team leads; develop successors and build resilience so operations don’t depend on any single individual. - Forecast staffing needs as the platform grows from initial deployment to full production scale. - **Establish the Production Operating Model** - Define the SRE operating model: staffing/coverage, escalation paths, on-call responsibilities, incident management, handoffs, production access, change management, and operational readiness requirements. - Create clear operational interfaces with Datacenter Operations, engineering teams, vendors, and other service owners. - Set and continuously improve production readiness standards, runbooks, operational procedures, failure-mode documentation, escalation processes, and incident response practices—prioritizing **automation and engineering over manual toil**. - Develop training, cross-training, simulations, and production incident-response exercises so the team can operate independently. - **Engineer Reliability Into the Platform** - Partner with platform engineering during development to ensure reliability, serviceability, observability, and operational requirements are built in before production. - Lead development of **SLOs**, operational health indicators, alerting standards, incident severity definitions, and reliability reporting for large-scale production infrastructure. - Build a culture that engineers out recurring operational problems via automation, improved observability, better platform design, and elimination of unnecessary toil. - Ensure the SRE organization can rapidly diagnose and mitigate issues across compute, networking, storage, and supporting infrastr

Listing freshness

CronJobs last confirmed this listing 46m ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord