Research Engineer / Performance Engineer, RL Distributed Systems
Anthropic · San Francisco, CA | New York City, NY | Seattle, WA
About this role
**Research Engineer / Performance Engineer — RL Distributed Systems** **About Anthropic** Anthropic’s mission is to create reliable, interpretable, and steerable AI systems—safe and beneficial for users and society. **About the role** Reinforcement learning is how Claude learns to reason, write code, and act autonomously over long horizons. At frontier scale, an RL run is a demanding distributed system: training, sampling, and environment execution run concurrently across large fleets of accelerators and hosts, continuously exchanging data while still making progress through hardware failures, load shifts, and changing research. As a **Research Engineer on the Distributed Systems team within RL Engineering**, you’ll work on the current bottleneck—e.g., scheduling/placement, data movement, sandboxed environment execution, storage & checkpointing, networking, fault tolerance, autoscaling, and observability. They’re looking for **generalists** who can move across layers, reason from first principles, and choose the most impactful problem. --- ## Key responsibilities - Design, build, and operate distributed systems that run RL at scale across training, sampling, and environment execution - Identify and remove system bottlenecks (scheduling, data movement, storage, networking, coordination) - Build fault tolerance end-to-end (failure detection, isolation, recovery) so long-running jobs keep progressing without human intervention - Design resource management and autoscaling so compute follows demand as run needs shift - Build observability to understand what a run is doing and why it slowed down, stalled, or produced unexpected results - Create automation to detect/remediate common problems and design safe interfaces for engineers and tools - Partner with researchers and performance engineers to preserve training correctness and avoid subtle nondeterminism - Remove failure classes through incident review, testing, and redesign; write clear design documents ## Minimum qualifications - Strong software engineering skills in **Python** and at least one systems language: **Rust, C++, or Go** - Experience designing/building/operating **large-scale distributed systems in production** - Deep understanding of distributed systems fundamentals (consistency, coordination, consensus, failure modes, recovery) - Ability to reason quantitatively about throughput, latency, and resource costs (compute, memory, storage, network) - Debugging experience for complex failures across many hosts/services (including failures not reproducible locally) - Strong written communication (design docs and incident writeups) ## Preferred qualifications - Experience running ML training or inference infrastructure at scale - Experience across multiple stack layers (scheduling, storage, networking, orchestration) - Experience building schedulers/autoscalers/resource management systems - Experience with **Kubernetes** and sandboxed/virtualized code execution at scale - Experience with high-performance networking (e.g., RDMA) or collective communication libraries - Experience with observability and automated remediation for large fleets - Experience with async Python frameworks (e.g., **Trio** or **asyncio**) - Familiarity with reinforcement learning or large language model training workloads --- ## Representative projects - Build a scheduler that places training/sampling/environment work across a heterogeneous cluster while respecting topology and failure domains - Creat
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.