Member of Technical Staff - Lead, Machines
Modal · New York
About this role
## About Modal AI needs a new infrastructure layer. At Modal, we’re building the systems that power modern AI workloads—enabling instant GPU access, sub-second container starts, and native storage for low-latency inference, fine-tuning, and production-ready sandboxes at scale. We serve customers including Lovable, Ramp, Cognition, DoorDash, and Suno. ## The Role (Member of Technical Staff – Lead, Machines) We’re looking for a strong technical lead to guide engineers designing, building, and maintaining the high-performance systems that make up our serverless platform. You’ll lead the team responsible for Modal’s **Machines layer**: - The fleet of **bare metal** and **cloud hosts** that run every **Function**, **Sandbox**, and **training job** - The **control plane** that provisions, images, monitors, and repairs machines ### What you’ll own - Full machine lifecycle: accepting and benchmarking new hardware from providers - Network bring-up - Kernel and image management - GPU and disk health tracking - Automated remediation of unhealthy hosts ### What you’ll do day-to-day - Manage a team of **3–8 engineers** - Stay hands-on across the stack, including: - BMCs, firmware, PXE, bootloaders - Linux networking, drivers - Distributed control-plane services - Shape the long-term technical direction further down the stack ## Requirements - **7+ years** of experience writing high-quality production code - **3+ years** of direct people management experience (project planning, growth, performance) - Experience operating **large fleets of physical hardware** at scale (bare metal provisioning, BMC/IPMI, PXE/network boot, firmware) **or** building control planes that manage them - Strong cloud skills - Strong knowledge of low-level OS foundations (Linux kernel, drivers, networking, file systems, containers, etc.) - Experience with hardware/colocation providers, including acceptance testing and benchmarking - Track record of setting technical direction and driving architectural decisions - Willingness to participate in **on-call** and respond to production incidents ## Nice to Haves - Experience with **GPUs** and the NVIDIA software stack (drivers, health monitoring, RDMA/NVLink) in production - Prior experience with **Go** ## What the Team is Working On - **Automatic remediation** of unhealthy machines (power cycling, reimaging, GPU recovery) to maximize uptime and reduce operator toil - **Automatic integration** of new CPU/GPU/storage servers while managing hardware and network heterogeneity - **Network health monitoring** and reliability across many datacenters; standardization of bare metal network configuration - **Automatic hardware acceptance testing and benchmarking** (CPU, disk, GPU, interconnect, network) - Custom **network bootloader**, **machine image pipeline**, and **kernel/firmware management** across the fleet
Listing freshness
CronJobs last confirmed this listing 7h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.