CronJobs

devops-sre jobs

Principal Datacenter Technologist

Graphcore · Milpitas, California, United States

hybridseniorPosted Oct 7, 2026LinuxAutomationTelemetryDiagnosticsData Center InfrastructureAI ComputingHigh-Performance Computing

Apply on the employer site

About this role

<h1>About Graphcore</h1> <p>Graphcore is a leading innovator in artificial intelligence computing. We develop hardware, software, and data center infrastructure that provide the specialized processing and systems capabilities needed to advance AI while improving the efficiency required for broad adoption.</p> <p>As part of SoftBank Group, Graphcore works alongside companies developing advanced technologies. Our teams bring together AI researchers, silicon designers, hardware and software engineers, and systems architects to solve complex technical problems across the computing stack.</p> <h1>The Opportunity</h1> <p>As Principal Datacenter Technologist, you will turn new platform architectures into repeatable operating models for Graphcore's AI data centers. You will lead the technical definition of how Graphcore introduces, deploys, services, and sustains new compute, networking, storage, memory, power, and cooling technologies at scale.</p> <p>Working within Advanced Architecture, you will connect platform design with data center engineering and site reliability operations. You will define new-product-introduction workflows, readiness criteria, diagnostics, and cross-functional handoffs so that new systems can move from architecture through deployment with clear ownership, supportability, and operational feedback.</p> <h1>What You Will Do</h1> <ul> <li>Own the technical operating model and new-product-introduction framework for new hardware and infrastructure technologies entering Graphcore data centers.</li> <li>Define end-to-end workflows for installation, configuration, validation, deployment, service, repair, upgrade, and sustained operation across the product lifecycle.</li> <li>Translate architecture requirements into data center readiness criteria covering software, firmware, compute, networking, storage, memory, rack integration, power, cooling, space, and site operations.</li> <li>Create clear runbooks, interface definitions, ownership models, acceptance criteria, and escalation paths that align architecture, engineering, data center, and site reliability teams.</li> <li>Lead operational readiness reviews and identify gaps in tooling, diagnostics, telemetry, serviceability, spares, documentation, and training before deployment.</li> <li>Serve as a senior escalation point for complex platform issues, ensuring that teams collect the right logs, telemetry, failure evidence, and environmental data to support root-cause analysis and disposition.</li> <li>Define and guide the deployment of automated diagnostic and telemetry systems that identify systemic hardware, software, quality, handling, and site-integration issues across the fleet.</li> <li>Establish quality and operational metrics for new technologies, analyze field trends, and feed findings back into architecture, design, supplier, and deployment decisions.</li> <li>Coordinate cross-functional resolution of issues spanning hardware, firmware, software, networking, storage, facilities, and site operations, with clear owners, decisions, and follow-through.</li> <li>Provide technical leadership, review implementation plans, mentor engineers, and communicate readiness, risks, tradeoffs, and recommendations to engineering and operations leaders.</li> </ul>...

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord