Member of Technical Staff - Storage
Modal · San Francisco
About this role
## About Modal AI needs a new infrastructure layer. At Modal, we’re building the storage and systems layer that enables low-latency inference, fine-tuning, and production-ready sandboxes at scale. Our customers include category-defining companies such as Lovable, Ramp, Cognition, DoorDash, and Suno. We provide instant GPU access, sub-second container starts, and native storage. ## The Role (Member of Technical Staff — Storage) You’ll design, build, and maintain the distributed object storage system that underpins every container image, volume, and checkpoint on Modal. ### What you’ll work on - **High-performance distributed storage & caching**: hundreds of petabytes of data replicated across multiple cloud object stores and a CDN, cached on local **NVMe** across a large fleet of workers in many datacenters, and shared **peer-to-peer** within each datacenter. - **Reduce cold-start latency**: make cold starts feel local even when data is hundreds of milliseconds away by designing caching, preloading, and peer-to-peer layers that hide object-store latency. - **Protect network and cost**: keep public ingress off saturated uplinks. - **Durability & cost at petabyte scale**: own streaming and batch replication between origins, and **garbage collection** over billions of objects. - **Cross-layer ownership**: work across local disk/page cache, distributed blob storage, garbage collection, and help shape what storage becomes next. ## Requirements - **5+ years** of experience writing high-quality production code - Experience building **high-performance distributed storage or caching systems** at large scale - Strong cloud skills, including deep familiarity with **object storage (S3 or similar)**, **CDNs**, and their consistency/throughput/cost characteristics - Strong knowledge of **low-level OS foundations** (Linux kernel, file systems, page cache, containers, etc.) - Willingness to participate in **on-call** and respond to production incidents ## Nice to Haves - Experience with **replication**, **content addressing**, and **consistency models** in multi-region or multi-cloud systems - Experience operating storage systems at scale (petabyte-scale datasets, high-throughput read/write paths, large-scale garbage collection or data migration) - Experience with **data engineering** at petabyte scale - Prior experience with **Rust** ## What the Team is Working On - **P2P sharing** of data across workers within a single datacenter to reduce ingress - Replicating data across multiple blob storage providers - Automating garbage collection across hundreds of petabytes of data - Deploying colocated storage clusters to datacenters to accelerate high-throughput customer workloads
Listing freshness
CronJobs last confirmed this listing 6h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.