CronJobs

devops-sre jobs

Member of Technical Staff — Compute Cluster

Causal · San Francisco

onsitestaffPosted Jul 19, 2026KubernetesSlurmDockerAWSGCPAzureCUDANCCL

Apply on the employer site

About this role

## Mission Our mission is general causal intelligence—AI that can (1) predict the future and (2) identify actions to alter it. We’re building a **Large Physics Foundation Model (LPM)** to learn verifiable cause-and-effect in physical systems (starting with **weather**). Everything we do—**training, evaluation, and serving**—runs on our **GPU fleet**. We’re looking for **Infrastructure Engineers** excited to tackle unsolved problems and build the supercomputing environment underneath it all. --- ## Role: Member of Technical Staff — Compute Cluster Design, build, and operate the compute platform that enables research to iterate rapidly at scale. --- ## Responsibilities - **Design, deploy, and operate** large distributed **GPU clusters** end-to-end: provisioning, imaging, upgrades, and capacity planning - **Extend scheduling and orchestration** systems (e.g., **Kubernetes, Slurm**) for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads - **Build cluster management software** that abstracts operations and provides a unified, self-serve interface for researchers and engineers - **Own cluster storage and artifact paths** for checkpoints and logs, including retention and lineage - **Monitor and improve reliability** continuously: error recovery and observability to catch failures before researchers do - **Partner with researchers** to unblock large-scale runs and advise on performance and placement trade-offs --- ## What We’re Looking For - Experience operating **large-scale GPU clusters** and **container orchestration** frameworks (e.g., **Kubernetes, Slurm, Docker**) - Strong systems background: **Linux, networking, storage, infrastructure-as-code** - Knowledge of cloud platforms (**GCP, AWS, or Azure**) and their ML/AI service offerings - Understanding of **monitoring, logging, observability**, and **version control** best practices for ML systems - Familiarity with **CUDA/NCCL** and performance profiling for distributed workloads - Ability to **own deliverables end-to-end**, from requirements through autonomous execution --- ## Background / Team Our founding team has built and deployed AI against the physical world in **robotics, drug discovery, and particle physics** at institutions including **DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN**.

Listing freshness

CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord