CronJobs

backend jobs

Performance Engineer, Inference Engine

Anthropic · San Francisco, CA | New York City, NY

hybridunknown$350,000–$350,000Posted Sep 9, 2026RustC++GPU/Accelerator programmingLLM inferenceTransformersRDMAPCIeHBM

Apply on the employer site

About this role

**Performance Engineer, Inference Engine** **About Anthropic** Anthropic’s mission is to create reliable, interpretable, and steerable AI systems—safe and beneficial for users and society. **About the Role** Anthropic’s inference engine is the software between accelerator kernels and the routing layer. It manages the full token path, including: - Batching requests - Laying the model out across chips - Managing memory for weights and activations - Coordinating every forward pass - Managing model state across requests This in-house system runs on all of Anthropic’s accelerator platforms, serving Claude to millions of users and powering research workloads. You’ll build and optimize this system at scale to improve **throughput, cost, reliability, and latency** across accelerator and cloud platforms. You’ll work with hardware and bandwidth constraints (e.g., FLOPs, HBM, PCIe, RDMA, network links), model where time/bytes go, and identify what sets the bound. The role is deeply technical and high-impact—ideal for engineers who enjoy accelerator programming, high-performance host/device coordination, and large-scale distributed systems. Familiarity with transformer architecture is a plus. **Recurring themes** - **Keep device utilization high:** avoid accelerators waiting on overheads. - **Reuse instead of recompute:** cache and reuse model state when cheaper than recomputation. - **Measure, model, then change:** build observability, model improvements, deploy, and iterate. - **Tokens you can trust:** maintain model quality across platforms and over time. - **Safety on every token:** support production safety systems with efficiency that doesn’t compromise robustness. **Minimum Qualifications** - A working mental model of LLM inference (prefill/decode on compute, memory, interconnect; what the host does meanwhile) - Proven ability to ramp quickly in unfamiliar systems and ship consequential changes - Strong systems programming (Rust, C++, or similar) with care for code quality and tests - Analytical performance mindset: profile → hypothesis → test → measure again - Low ego: ask naive questions, take feedback well, and pick up slack - Enjoy pair programming and care about the societal impacts of your work **Preferred Qualifications** - Experience inside an LLM serving engine and awareness of where abstractions strain - GPU/accelerator programming - OS internals - Language modeling with transformers - Experience building an allocator, cache, scheduler, or high-bandwidth transport - Fluency in Rust - Experience making systems reproducible (determinism, replay, property-based tests) **Compensation (Annual Salary)** - **$350,000 — $850,000 USD** **Logistics** - **Minimum education:** Bachelor’s degree (or equivalent combination of education, training, and/or experience) - **Required field of study:** Field relevant to the role (via coursework, training, or professional experience) - **Minimum years of experience:** Correlates with internal job level requirements - **Location-based hybrid policy:** Expect to be in one of Anthropic’s offices at least **25%** of the time (some roles may require more) - **Visa sponsorship:** Anthropic sponsors visas, but not for every role/candidate. If offered, they’ll make reasonable efforts and retain an immigration lawyer to help. **How we’re different** Anthropic emphasizes big-science, highly collaborative work on a few large-scale research efforts with an emphasis on impact and communication. **Come wo

Listing freshness

CronJobs last confirmed this listing 17h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord