Member of Technical Staff — Training Infrastructure
Causal · San Francisco
About this role
## Member of Technical Staff — Training Infrastructure ### About the role Our mission is general causal intelligence—AI that can (1) predict the future and (2) identify the actions to alter it. To reach this breakthrough, we’re building a **Large Physics Foundation Model (LPM)**, leveraging the fact that physical systems are governed by verifiable cause and effect. You’ll help train these models by building **large-scale training infrastructure** that is **fast, efficient, and reliable**, so every GPU cycle accelerates research progress. ### Responsibilities - Design, implement, and optimize **distributed training systems** that scale across **thousands of GPUs** - Research and test **parallelization strategies** and **numerical precision trade-offs** across model scales (including architectures that don’t map cleanly onto existing LLM training stacks) - **Analyze, profile, and debug** low-level GPU operations to maximize throughput and hardware utilization - Build reusable frameworks for **checkpointing, fault tolerance, and reproducibility** that remain robust under rapid research iteration - Collaborate with researchers to bring novel model architectures from **prototype to full scale** - Stay up-to-date on research and bring new ideas into production training workflows ### What we’re looking for We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains. - Demonstrated proficiency with **distributed training frameworks/techniques** (e.g., **FSDP, DeepSpeed, Megatron, PyTorch, JAX/XLA**) for training large foundation models - Strong grasp of state-of-the-art techniques for optimizing training workloads: - parallelism strategies - memory optimization - mixed precision - communication overlap - Ability to **profile and debug performance** in complex codebases, from framework internals down to kernels and collectives - Deep understanding of deep learning frameworks (e.g., **PyTorch, JAX**) and their underlying system architectures **Bonus** - Contributions to open-source ML infrastructure (e.g., **PyTorch, Megatron-LM, DeepSpeed, XLA**)
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.