Staff + Sr. Software Engineer, Scaling
Anthropic · San Francisco, CA | Seattle, WA
About this role
## About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems—safe and beneficial for users and society. ## About the role The Inference team builds and scales the critical systems that serve Claude to millions of users worldwide. You’ll help deliver Claude by serving models through large compute-agnostic inference deployments, owning the full stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team has a dual mandate: - **Maximize compute efficiency** to reliably serve explosive customer growth - **Enable breakthrough research** by providing high-performance inference infrastructure for next-generation models Inference systems are highly performance-sensitive distributed systems, requiring sophisticated routing, scaling, and networking across a large fleet. ## Key responsibilities - Design, build, and maintain distributed systems that serve Claude to millions of users worldwide - Develop resilient, flexible systems that adapt in real time to real-world events - Build intelligent request routing, load balancing, and traffic management across thousands of accelerators and multiple cloud providers - Maximize compute efficiency and optimize cost via autoscaling and orchestration across production, research, and experimental workloads - Build and operate production-grade deployment pipelines for releasing new models to users - Provide high-performance inference infrastructure that enables researchers to develop next-generation models - Integrate new AI accelerator platforms and support inference for new model architectures ## Minimum qualifications - Significant software engineering experience, particularly with distributed systems - Results-oriented, with a bias towards flexibility and impact - Willingness to pick up slack, even if it goes outside your job description - Desire to learn more about machine learning systems and infrastructure - Thrive in environments where technical excellence drives both business results and research breakthroughs - Care about the societal impacts of your work ## Preferred qualifications - Experience with high-performance, large-scale distributed systems - Experience implementing and deploying machine learning systems at scale - Experience with load balancing, request routing, or traffic management systems - Familiarity with LLM inference optimization, batching, and caching strategies - Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure) - Proficiency in Python or Rust ## Representative projects - Designing intelligent routing algorithms to optimize request distribution across many accelerators in different environments - Autoscaling the compute fleet to match supply with demand across production, research, and experimental workloads - Building production-grade deployment pipelines for releasing new models to millions of users reliably - Contributing to new inference features - Supporting inference for new model architectures - Analyzing observability data to tune performance based on real-world production workloads - Managing multi-region deployments and geographic routing for global customers ## Compensation (Annual Salary) **$320,000 — $485,000 USD** ## Logistics - **Minimum education:** Bachelor’s degree or equivalent combination of education, training, and/or experience - **Required field of study:** Field relevant to the role (via coursework, training, or professional experience) - **Mi
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.