CronJobs

devops-sre jobs

Member of Technical Staff — Training Infrastructure

Causal · San Francisco

onsitestaffPosted Oct 29, 2025PyTorchJAXXLAFSDPDeepSpeedMegatron-LMdistributed trainingmixed precision

Apply on the employer site

About this role

## Member of Technical Staff — Training Infrastructure ### About the role Our mission is general causal intelligence—AI that can (1) predict the future and (2) identify the actions to alter it. To reach this breakthrough, we’re building a **Large Physics Foundation Model (LPM)**, leveraging the fact that physical systems are governed by verifiable cause and effect. You’ll help train these models by building **large-scale training infrastructure** that is **fast, efficient, and reliable**, so every GPU cycle accelerates research progress. ### Responsibilities - Design, implement, and optimize **distributed training systems** that scale across **thousands of GPUs** - Research and test **parallelization strategies** and **numerical precision trade-offs** across model scales (including architectures that don’t map cleanly onto existing LLM training stacks) - **Analyze, profile, and debug** low-level GPU operations to maximize throughput and hardware utilization - Build reusable frameworks for **checkpointing, fault tolerance, and reproducibility** that remain robust under rapid research iteration - Collaborate with researchers to bring novel model architectures from **prototype to full scale** - Stay up-to-date on research and bring new ideas into production training workflows ### What we’re looking for We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains. - Demonstrated proficiency with **distributed training frameworks/techniques** (e.g., **FSDP, DeepSpeed, Megatron, PyTorch, JAX/XLA**) for training large foundation models - Strong grasp of state-of-the-art techniques for optimizing training workloads: - parallelism strategies - memory optimization - mixed precision - communication overlap - Ability to **profile and debug performance** in complex codebases, from framework internals down to kernels and collectives - Deep understanding of deep learning frameworks (e.g., **PyTorch, JAX**) and their underlying system architectures **Bonus** - Contributions to open-source ML infrastructure (e.g., **PyTorch, Megatron-LM, DeepSpeed, XLA**)

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord