Member of Technical Staff - Lead, Storage
Modal · San Francisco
About this role
## About Modal AI needs a new infrastructure layer. At Modal, we’re building the layer underneath modern AI workloads—so customers can get instant GPU access, sub-second container starts, and native storage for low-latency inference, fine-tuning, and production-ready sandboxes at scale. Modal serves category-defining companies including Lovable, Ramp, Cognition, DoorDash, and Suno. ## The Role — Member of Technical Staff, Lead (Storage) We’re looking for a strong technical lead to guide engineers designing, building, and maintaining the novel, high-performance systems that power Modal’s serverless platform. You’ll lead the team responsible for the **distributed object storage system** that underpins every container image, volume, and checkpoint on Modal: - **Hundreds of petabytes** of data - Replicated across **multiple cloud object stores** and a **CDN** - Cached on **local NVMe** across a large fleet of workers in many datacenters - Shared **peer-to-peer** within each datacenter ### What you’ll do - Set technical direction for storage primitives that other teams build on (filesystems, training, sandboxes) - Balance **durability, latency, throughput, and cost** - Own the roadmap from today’s hardest problems to major architectural bets, including: - Garbage collection at petabyte scale - Active-active replication - Rate limiting that protects upstream systems without wasting utilization - Storage colocated with GPUs - Tiered writes - Capacity planning against provider limits - Manage a team of **3–8 engineers** while staying hands-on across the stack: - Local disk and page cache - Distributed blob storage and garbage collection - Drive **observability, automation, and on-call practices** to keep the system healthy as it scales ## Requirements - **7+ years** of experience writing high-quality production code - **3+ years** of direct people management experience (leading engineers through planning, growth, and performance) - Experience building **high-performance distributed storage or caching** systems at large scale - Strong cloud skills, including deep familiarity with **object storage (S3 or similar)**, **CDNs**, and their consistency/throughput/cost characteristics - Strong knowledge of **low-level OS foundations** (Linux kernel, file systems, page cache, containers, etc.) - Experience with **replication, content addressing, and consistency models** in multi-region/multi-cloud systems - Experience operating storage systems at scale (petabyte-scale datasets, high-throughput read/write paths, large-scale garbage collection or data migration), including owning **cost and capacity planning** - Track record of setting technical direction and driving architectural decisions across a team - Willingness to participate in **on-call** and respond to production incidents ## Nice to Haves - Experience with **data engineering** at petabyte scale - Prior experience with **Rust** ## Key Areas the Team Is Working On - **P2P sharing** of data across workers within a single datacenter to reduce ingress - Replicating data across **multiple blob storage providers** - Automating **garbage collection** across hundreds of petabytes - Deploying **colocated storage clusters** to datacenters to accelerate high-throughput customer workloads
Listing freshness
CronJobs last confirmed this listing 4h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.