Member of Technical Staff — Compute Cluster
Causal · San Francisco
About this role
## Mission Our mission is general causal intelligence—AI that can (1) predict the future and (2) identify actions to alter it. We’re building a **Large Physics Foundation Model (LPM)** to learn verifiable cause-and-effect in physical systems (starting with **weather**). Everything we do—**training, evaluation, and serving**—runs on our **GPU fleet**. We’re looking for **Infrastructure Engineers** excited to tackle unsolved problems and build the supercomputing environment underneath it all. --- ## Role: Member of Technical Staff — Compute Cluster Design, build, and operate the compute platform that enables research to iterate rapidly at scale. --- ## Responsibilities - **Design, deploy, and operate** large distributed **GPU clusters** end-to-end: provisioning, imaging, upgrades, and capacity planning - **Extend scheduling and orchestration** systems (e.g., **Kubernetes, Slurm**) for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads - **Build cluster management software** that abstracts operations and provides a unified, self-serve interface for researchers and engineers - **Own cluster storage and artifact paths** for checkpoints and logs, including retention and lineage - **Monitor and improve reliability** continuously: error recovery and observability to catch failures before researchers do - **Partner with researchers** to unblock large-scale runs and advise on performance and placement trade-offs --- ## What We’re Looking For - Experience operating **large-scale GPU clusters** and **container orchestration** frameworks (e.g., **Kubernetes, Slurm, Docker**) - Strong systems background: **Linux, networking, storage, infrastructure-as-code** - Knowledge of cloud platforms (**GCP, AWS, or Azure**) and their ML/AI service offerings - Understanding of **monitoring, logging, observability**, and **version control** best practices for ML systems - Familiarity with **CUDA/NCCL** and performance profiling for distributed workloads - Ability to **own deliverables end-to-end**, from requirements through autonomous execution --- ## Background / Team Our founding team has built and deployed AI against the physical world in **robotics, drug discovery, and particle physics** at institutions including **DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN**.
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.