CronJobs

devops-sre jobs

Member of Technical Staff — Inference Infrastructure

Causal · San Francisco

onsiteseniorPosted Jul 19, 2026PyTorchJAXKubernetesRaySlurmTensorRTvLLMTriton

Apply on the employer site

About this role

## Member of Technical Staff — Inference Infrastructure ### Mission Our mission is general causal intelligence: AI that can **(1) predict the future** and **(2) identify actions to alter it**. To achieve this, we’re building a **Large Physics Foundation Model (LPM)**—because physical systems are governed by **verifiable cause and effect**. We believe scaling on physics will enable the understanding of causality needed to **predict and control physical systems**, starting with **weather**. Our founding team has built and deployed AI against the physical world in **robotics, drug discovery, and particle physics** (e.g., DeepMind, Waymo, Cruise, Insitro, Nabla Bio, CERN). ### Why this role Progress on an LPM is gated by how fast we can evaluate it: **large-scale backtesting** against decades of physical observations, **ensemble generation**, and **rollout evaluation** across model scales. **Your mission:** make inference **so fast and cheap** that evaluation never gates research. --- ### Responsibilities - Build **high-throughput inference** systems for large-scale evaluation, backtesting, and scoring against historical physical observations - Design and implement techniques that improve **latency, throughput, and efficiency** for real-time inference - Optimize the inference stack to fully utilize **hardware FLOPs, bandwidth, and memory** - Extend orchestration frameworks (e.g., **Kubernetes, Ray, Slurm**) for **distributed inference** and large-batch evaluation sweeps - Establish standards for **reliability, observability, and reproducibility** across the inference stack, so every evaluation is trustworthy and repeatable - Collaborate with researchers to enable high-performance inference for **novel architectures** as they emerge --- ### What we’re looking for We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains. - Experience building or optimizing inference and serving systems for **throughput and latency** (e.g., **TensorRT**) - Understanding of **distributed compute**, **GPU parallelism**, and **hardware-aware optimization** - Deep familiarity with deep learning frameworks (e.g., **PyTorch, JAX**) and their underlying system architectures - Strong engineering skills: performant, maintainable code and the ability to debug complex codebases - **Bonus:** contributions to open-source inference or systems infrastructure (e.g., **vLLM, SGLang, Triton**)

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord