Performance Engineer, Inference Engine
Anthropic · San Francisco, CA | New York City, NY
About this role
**Performance Engineer, Inference Engine** **About Anthropic** Anthropic’s mission is to create reliable, interpretable, and steerable AI systems—safe and beneficial for users and society. **About the Role** Anthropic’s inference engine is the software between accelerator kernels and the routing layer. It manages the full token path, including: - Batching requests - Laying the model out across chips - Managing memory for weights and activations - Coordinating every forward pass - Managing model state across requests This in-house system runs on all of Anthropic’s accelerator platforms, serving Claude to millions of users and powering research workloads. You’ll build and optimize this system at scale to improve **throughput, cost, reliability, and latency** across accelerator and cloud platforms. You’ll work with hardware and bandwidth constraints (e.g., FLOPs, HBM, PCIe, RDMA, network links), model where time/bytes go, and identify what sets the bound. The role is deeply technical and high-impact—ideal for engineers who enjoy accelerator programming, high-performance host/device coordination, and large-scale distributed systems. Familiarity with transformer architecture is a plus. **Recurring themes** - **Keep device utilization high:** avoid accelerators waiting on overheads. - **Reuse instead of recompute:** cache and reuse model state when cheaper than recomputation. - **Measure, model, then change:** build observability, model improvements, deploy, and iterate. - **Tokens you can trust:** maintain model quality across platforms and over time. - **Safety on every token:** support production safety systems with efficiency that doesn’t compromise robustness. **Minimum Qualifications** - A working mental model of LLM inference (prefill/decode on compute, memory, interconnect; what the host does meanwhile) - Proven ability to ramp quickly in unfamiliar systems and ship consequential changes - Strong systems programming (Rust, C++, or similar) with care for code quality and tests - Analytical performance mindset: profile → hypothesis → test → measure again - Low ego: ask naive questions, take feedback well, and pick up slack - Enjoy pair programming and care about the societal impacts of your work **Preferred Qualifications** - Experience inside an LLM serving engine and awareness of where abstractions strain - GPU/accelerator programming - OS internals - Language modeling with transformers - Experience building an allocator, cache, scheduler, or high-bandwidth transport - Fluency in Rust - Experience making systems reproducible (determinism, replay, property-based tests) **Compensation (Annual Salary)** - **$350,000 — $850,000 USD** **Logistics** - **Minimum education:** Bachelor’s degree (or equivalent combination of education, training, and/or experience) - **Required field of study:** Field relevant to the role (via coursework, training, or professional experience) - **Minimum years of experience:** Correlates with internal job level requirements - **Location-based hybrid policy:** Expect to be in one of Anthropic’s offices at least **25%** of the time (some roles may require more) - **Visa sponsorship:** Anthropic sponsors visas, but not for every role/candidate. If offered, they’ll make reasonable efforts and retain an immigration lawyer to help. **How we’re different** Anthropic emphasizes big-science, highly collaborative work on a few large-scale research efforts with an emphasis on impact and communication. **Come wo
Listing freshness
CronJobs last confirmed this listing 17h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.