CronJobs

devops-sre jobs

Senior Infrastructure Software Engineer

Lightning AI · New York, New York, United States

hybridsenior$180,000–$180,000Posted Aug 31, 2026PythonLinuxKubernetesAPIsBare-metal provisioningGPU infrastructureHPCAutomation

Apply on the employer site

About this role

**Senior Infrastructure Software Engineer** **About Lightning AI** Lightning AI is the company behind PyTorch Lightning, building an end-to-end platform for developing, training, and deploying AI systems. Through our merger with Voltage Park, we combine developer-first software with cost-efficient, large-scale compute. We serve solo researchers, startups, and large enterprises globally. **The Role** Join our Infrastructure Engineering team to build production software systems that operate and manage Lightning AI's large-scale GPU and bare-metal infrastructure. You'll design services, APIs, tooling, and automation that bridge physical infrastructure and customer-facing systems for AI/ML training, inference, and HPC workloads. **What You'll Do** • Design and operate production services for managing large-scale bare-metal and GPU infrastructure • Build systems for server discovery, provisioning, configuration, and lifecycle management • Develop reliable automation that reduces manual work and improves consistency at scale • Build telemetry, logging, and observability capabilities for infrastructure health • Partner with Network, Infrastructure Operations, and Platform Engineering teams • Make pragmatic architecture decisions balancing reliability, scalability, and speed **Required Qualifications** • 8+ years of professional software engineering or infrastructure engineering experience • Strong software fundamentals and production backend systems experience (Python or similar) • Strong Linux production environment experience • Experience building APIs, tooling, or automation for infrastructure at scale • Familiarity with containerization and orchestration concepts • Understanding of HPC and bare-metal infrastructure fundamentals • Experience in fast-paced startup environments with high ownership **Ideal Experience** • Bare-metal hardware troubleshooting and provisioning (PXE/iPXE, BMC, Redfish, IPMI) • GPU server experience • Network infrastructure experience (SONiC, Palo Alto, Juniper) • High-performance storage systems (VAST) • AI/ML or HPC infrastructure at scale **Compensation & Benefits** 💰 $180,000 – $220,000 USD base salary 📊 Discretionary bonus and meaningful equity 🏥 Comprehensive health coverage (medical, dental, vision) 💼 401(k) matching (U.S.) and pension contributions (U.K.) ⏰ Flexible time off **Location** Based in NYC, SF, Seattle, or London with minimum 2 in-office days/week. No visa sponsorship available at this time.

Listing freshness

CronJobs last confirmed this listing 2d ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord