Member of Technical Staff - Machines
Modal · San Francisco
About this role
## About Modal AI needs a new infrastructure layer. At Modal, we’re building the systems that power modern AI workloads—enabling instant GPU access, sub-second container starts, and native storage for low-latency inference, fine-tuning, and production-ready sandboxes at scale. Modal serves customers including Lovable, Ramp, Cognition, DoorDash, and Suno. ## The Role — Member of Technical Staff (Machines) We’re looking for strong engineers to design, build, and maintain the novel, high-performance systems behind Modal’s serverless platform. You’ll work on **Modal’s Machines layer**, including: - The fleet of **bare metal and cloud hosts** that run every Function, Sandbox, and training job - The **control plane** that provisions, images, monitors, and repairs machines ### What you’ll do - **Automate integration of new capacity** from a growing set of hardware providers - Audit and benchmark hosts and clusters - Maintain machine images - Configure **GPUs, RDMA, networking, and storage** - Get machines into production - Build automation to keep the fleet healthy **without human intervention** - Detect bad GPUs, thermals, and disks - Debug across the stack between hardware and software - Kernel panics, boot-time broadcast storms - Enabling container runtime support for new architectures/platforms ## Requirements - **5+ years** of experience writing high-quality production code - Experience operating fleets of physical hardware (e.g., bare metal provisioning, BMC/IPMI, PXE/network boot, firmware) **or** building control planes to manage them - Strong cloud skills - Strong knowledge of low-level OS foundations (Linux kernel, drivers, networking, file systems, containers, etc.) - Ability to debug across layers (e.g., BGP issues, Linux RPS, vBIOS bugs, Python control-plane services) - Willingness to participate in **on-call** and respond to production incidents ## Nice to Haves - Production experience with **GPUs / NVIDIA software stack** (drivers, health monitoring, XIDs, RDMA/NVLink) - Prior experience with **Go** ## What the Team is Working On - Automatic remediation of unhealthy machines (power cycling, reimaging, GPU recovery) to maximize uptime and reduce operator toil - Automatic integration of new CPU/GPU/storage servers while handling hardware and network heterogeneity - Network health monitoring and reliability across many datacenters; standardizing bare metal network configuration - Automatic hardware acceptance testing and benchmarking (CPU, disk, GPU, interconnect, network) - Custom network bootloader, machine image pipeline, and kernel/firmware management across the fleet
Listing freshness
CronJobs last confirmed this listing 4h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.