Staff+ Site Reliability Engineer, Safeguards ML Infra
Anthropic · Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY
About this role
## About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems—AI that is safe and beneficial for users and society. ## About the role The **Safeguards ML Infra** team designs, builds, and operates the production infrastructure that powers Claude’s safety systems. You’ll own critical backend services on the token generation path and the operational work required to get safeguards safely into production. This includes configuring, verifying, and rolling out safeguards across every platform Claude runs on (1P, AWS Bedrock, GCP Vertex, etc.), and leading incident response when issues arise. In this role, you’ll ensure safeguards are properly configured and deployed for model launches, and you’ll own off-cycle deployment of new safety classifiers—canarying changes, verifying that the right safeguards are provably live on the right models, and holding rollback authority when something looks wrong. You’ll also work to reduce manual checklists by turning runbooks into tooling and one-off deploys into repeatable pipelines. ## What you’ll do - **Launch captain for model releases:** stand up, configure, and verify safeguards for every new model; serve as the safeguards point of contact during release windows. - **Own off-cycle classifier deployments:** canary rollouts, run post-deploy validations, and investigate discrepancies. - **Cross-platform verification:** ensure the correct safeguards are provably live on the correct models across all deployment platforms, and detect/eliminate configuration drift. - **Automate away prior work:** convert launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. - **Use Claude aggressively** to enable safe agentic operations of safety-critical systems. ## What we’re looking for Engineers with deep experience in **production change management at scale**—people who have owned deploy pipelines, config management systems, rollout safety, or launch readiness under real production pressure. Familiarity with ML research or transformer architectures is **not required**; you’ll learn on the job. We prioritize **production judgment**: a track record of shipping changes to critical systems safely and automating yourself out of the work you did last quarter.
Listing freshness
CronJobs last confirmed this listing 1d ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.