CronJobs

devops-sre jobs

Staff+ Site Reliability Engineer, Safeguards ML Infra

Anthropic · Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY

unknownstaff$405,000–$405,000Posted Aug 11, 2026AWSGCPMLdeployment pipelinesconfig management

Apply on the employer site

About this role

## About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems—AI that is safe and beneficial for users and society. ## About the role The **Safeguards ML Infra** team designs, builds, and operates the production infrastructure that powers Claude’s safety systems. You’ll own critical backend services on the token generation path and the operational work required to get safeguards safely into production. This includes configuring, verifying, and rolling out safeguards across every platform Claude runs on (1P, AWS Bedrock, GCP Vertex, etc.), and leading incident response when issues arise. In this role, you’ll ensure safeguards are properly configured and deployed for model launches, and you’ll own off-cycle deployment of new safety classifiers—canarying changes, verifying that the right safeguards are provably live on the right models, and holding rollback authority when something looks wrong. You’ll also work to reduce manual checklists by turning runbooks into tooling and one-off deploys into repeatable pipelines. ## What you’ll do - **Launch captain for model releases:** stand up, configure, and verify safeguards for every new model; serve as the safeguards point of contact during release windows. - **Own off-cycle classifier deployments:** canary rollouts, run post-deploy validations, and investigate discrepancies. - **Cross-platform verification:** ensure the correct safeguards are provably live on the correct models across all deployment platforms, and detect/eliminate configuration drift. - **Automate away prior work:** convert launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. - **Use Claude aggressively** to enable safe agentic operations of safety-critical systems. ## What we’re looking for Engineers with deep experience in **production change management at scale**—people who have owned deploy pipelines, config management systems, rollout safety, or launch readiness under real production pressure. Familiarity with ML research or transformer architectures is **not required**; you’ll learn on the job. We prioritize **production judgment**: a track record of shipping changes to critical systems safely and automating yourself out of the work you did last quarter.

Listing freshness

CronJobs last confirmed this listing 1d ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord