CronJobs

backend jobs

Research Engineer / Performance Engineer, RL Distributed Systems

Anthropic · San Francisco, CA | New York City, NY | Seattle, WA

hybridunknown$500,000–$500,000Posted Sep 29, 2026PythonRustGoKubernetesasyncioRDMA

Apply on the employer site

About this role

**Research Engineer / Performance Engineer — RL Distributed Systems** **About Anthropic** Anthropic’s mission is to create reliable, interpretable, and steerable AI systems—safe and beneficial for users and society. **About the role** Reinforcement learning is how Claude learns to reason, write code, and act autonomously over long horizons. At frontier scale, an RL run is a demanding distributed system: training, sampling, and environment execution run concurrently across large fleets of accelerators and hosts, continuously exchanging data while still making progress through hardware failures, load shifts, and changing research. As a **Research Engineer on the Distributed Systems team within RL Engineering**, you’ll work on the current bottleneck—e.g., scheduling/placement, data movement, sandboxed environment execution, storage & checkpointing, networking, fault tolerance, autoscaling, and observability. They’re looking for **generalists** who can move across layers, reason from first principles, and choose the most impactful problem. --- ## Key responsibilities - Design, build, and operate distributed systems that run RL at scale across training, sampling, and environment execution - Identify and remove system bottlenecks (scheduling, data movement, storage, networking, coordination) - Build fault tolerance end-to-end (failure detection, isolation, recovery) so long-running jobs keep progressing without human intervention - Design resource management and autoscaling so compute follows demand as run needs shift - Build observability to understand what a run is doing and why it slowed down, stalled, or produced unexpected results - Create automation to detect/remediate common problems and design safe interfaces for engineers and tools - Partner with researchers and performance engineers to preserve training correctness and avoid subtle nondeterminism - Remove failure classes through incident review, testing, and redesign; write clear design documents ## Minimum qualifications - Strong software engineering skills in **Python** and at least one systems language: **Rust, C++, or Go** - Experience designing/building/operating **large-scale distributed systems in production** - Deep understanding of distributed systems fundamentals (consistency, coordination, consensus, failure modes, recovery) - Ability to reason quantitatively about throughput, latency, and resource costs (compute, memory, storage, network) - Debugging experience for complex failures across many hosts/services (including failures not reproducible locally) - Strong written communication (design docs and incident writeups) ## Preferred qualifications - Experience running ML training or inference infrastructure at scale - Experience across multiple stack layers (scheduling, storage, networking, orchestration) - Experience building schedulers/autoscalers/resource management systems - Experience with **Kubernetes** and sandboxed/virtualized code execution at scale - Experience with high-performance networking (e.g., RDMA) or collective communication libraries - Experience with observability and automated remediation for large fleets - Experience with async Python frameworks (e.g., **Trio** or **asyncio**) - Familiarity with reinforcement learning or large language model training workloads --- ## Representative projects - Build a scheduler that places training/sampling/environment work across a heterogeneous cluster while respecting topology and failure domains - Creat

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord