Engineering Manager, Kubernetes Infrastructure (Bare Metal)
CoreWeave · Livingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA
About this role
**CoreWeave — Engineering Manager, Kubernetes Infrastructure (Bare Metal)** ## What You’ll Do CoreWeave is looking for an **Engineering Manager** to lead a team building and operating **Kubernetes infrastructure on bare metal**. This team is close to the core of the platform and is responsible for the **reliability, scalability, and operational excellence** of systems powering high-performance **AI/ML workloads**. You’ll lead engineers across **cluster lifecycle**, **platform reliability**, **infrastructure automation**, and the operational systems that help Kubernetes run **predictably at scale** on dedicated hardware. ## In this role, you will: - Lead a team of engineers responsible for **Kubernetes infrastructure on bare metal** - Set goals, priorities, and execution plans to ensure **reliable delivery** - Partner with senior ICs and adjacent teams on the roadmap for **upgrades, reliability, observability, and infrastructure automation** - Improve operational excellence across **incident response, on-call health, root-cause analysis, and service ownership** - Drive best practices for **safe change management, testing, rollout quality, and production readiness** - Support design/operation of platform capabilities for **provisioning, patching, upgrades, scaling, and troubleshooting** - Build strong cross-functional relationships with **compute, networking, storage, security, and product** - Hire, coach, and develop engineers; build a **high-accountability, high-trust** culture - Establish mechanisms for **planning, prioritization, execution tracking, and continuous improvement** - Translate complex platform/infrastructure work into **clear business and customer value** ## Who You Are - Experience managing an **infrastructure/platform/SRE-oriented** engineering team - Strong technical depth in **Kubernetes, distributed systems, and production infrastructure** - Experience operating Kubernetes in complex environments (ideally **bare metal**, hybrid, or performance-sensitive systems) - Familiarity with **cluster lifecycle management** (provisioning, upgrades, node operations, observability, reliability) - Track record improving **execution, engineering quality, and operational maturity** - Experience leading **incident response** cultures and driving reliability improvements - Strong partnership skills across engineering, product, and operations - Ability to coach engineers and create clarity in ambiguous/fast-scaling environments - Strong written/verbal communication and ability to explain **technical trade-offs** ## Preferred Experience - Experience in **GPU-heavy, HPC, or ML infrastructure** environments - Experience with **bare-metal infrastructure**, server lifecycle operations, or low-level troubleshooting - Familiarity with Kubernetes **networking, storage, and security** primitives in production - Experience building internal platform products used by other teams or customers - Experience with infrastructure automation using tools such as **Go, Python, controllers/operators, or configuration management** ## What success looks like **First 90 days** - Build trust with the team and key partners - Assess team health, role clarity, roadmap risks, and operational pain points - Establish/tighten operating rhythms for planning, execution, and incident follow-up - Create a clear view of highest-value reliability/scalability opportunities **First 6 months** - Improve predictability of execution and service ownership - Raise the qual
Listing freshness
CronJobs last confirmed this listing 11h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.