Member of Technical Staff — Data Infrastructure
Causal · San Francisco
About this role
## Member of Technical Staff — Data Infrastructure ### Mission We’re building a **Large Physics foundation Model (LPM)** to enable AI that can **predict the future** and **identify actions to alter it**. Physical systems offer verifiable cause-and-effect, and we’re starting with domains like **weather**. Your work will power the **data platform** beneath everything—making datasets **cheap to ingest**, **fast to query**, and **immediately available** for training. ### Responsibilities - **Design and operate petabyte-scale storage**: lakehouse architecture, file formats, and data layout optimized for both **batch** and **real-time** queries. - **Own shared compute and orchestration** (e.g., **Spark**, **Ray**, workflow scheduling) for ingestion and research pipelines. - **Optimize the end-to-end data strategy** from storage to loading, including **high-throughput data loading** into training up to the **tensor boundary**. - Build systems for **cataloging, deduplication, lineage, search, and reproducibility** across the full data lifecycle. - Implement **platform-level quality and monitoring** tooling for data and research teams. - **Scale infrastructure** to improve engineering velocity while ensuring **reliability**, with monitoring and alerting. - Work across the full data lifecycle, including building and operating **ingestion pipelines** for critical data sources. ### What We’re Looking For - A relentless approach to problem-solving, **rapid execution**, and the ability to quickly learn in unfamiliar domains. - Demonstrated experience building **large-scale data pipelines** and **distributed compute** systems (e.g., **Spark**, **Ray**, **Beam**). - Knowledge of state-of-the-art tools and methods for **data ingestion, storage, and loading**, including how file formats and storage systems impact performance and scalability (e.g., **Parquet**, **Zarr**, **Delta Lake**). - Deep familiarity with **cloud infrastructure**, **data lake architectures**, and **batch + streaming** pipelines. - Understanding of how **data loading throughput** affects large-scale training, with experience optimizing it. - Ability to own deliverables **end-to-end**, from translating requirements to autonomously driving execution. ### Background / Team Our founding team has built and deployed AI against the physical world in areas including **robotics, drug discovery, and particle physics**, with experience from organizations such as **DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN**.
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.