Member of Technical Staff — Inference Infrastructure
Causal · San Francisco
About this role
## Member of Technical Staff — Inference Infrastructure ### Mission Our mission is general causal intelligence: AI that can **(1) predict the future** and **(2) identify actions to alter it**. To achieve this, we’re building a **Large Physics Foundation Model (LPM)**—because physical systems are governed by **verifiable cause and effect**. We believe scaling on physics will enable the understanding of causality needed to **predict and control physical systems**, starting with **weather**. Our founding team has built and deployed AI against the physical world in **robotics, drug discovery, and particle physics** (e.g., DeepMind, Waymo, Cruise, Insitro, Nabla Bio, CERN). ### Why this role Progress on an LPM is gated by how fast we can evaluate it: **large-scale backtesting** against decades of physical observations, **ensemble generation**, and **rollout evaluation** across model scales. **Your mission:** make inference **so fast and cheap** that evaluation never gates research. --- ### Responsibilities - Build **high-throughput inference** systems for large-scale evaluation, backtesting, and scoring against historical physical observations - Design and implement techniques that improve **latency, throughput, and efficiency** for real-time inference - Optimize the inference stack to fully utilize **hardware FLOPs, bandwidth, and memory** - Extend orchestration frameworks (e.g., **Kubernetes, Ray, Slurm**) for **distributed inference** and large-batch evaluation sweeps - Establish standards for **reliability, observability, and reproducibility** across the inference stack, so every evaluation is trustworthy and repeatable - Collaborate with researchers to enable high-performance inference for **novel architectures** as they emerge --- ### What we’re looking for We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains. - Experience building or optimizing inference and serving systems for **throughput and latency** (e.g., **TensorRT**) - Understanding of **distributed compute**, **GPU parallelism**, and **hardware-aware optimization** - Deep familiarity with deep learning frameworks (e.g., **PyTorch, JAX**) and their underlying system architectures - Strong engineering skills: performant, maintainable code and the ability to debug complex codebases - **Bonus:** contributions to open-source inference or systems infrastructure (e.g., **vLLM, SGLang, Triton**)
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.