Staff Software Engineer, MetalDev
CoreWeave · New York, NY / Sunnyvale, CA
About this role
**Staff Software Engineer — MetalDev (CoreWeave)** CoreWeave is *The Essential Cloud for AI™*. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. --- ### About the Role As a **Staff Software Engineer** in the **Compute Architecture** organization, you’ll help build the software systems that operate the backbone of large-scale GPU data centers. The **METALDEV** team builds **Go-based distributed services** to bring new infrastructure online, manage hardware lifecycle workflows, monitor production health, and automate safe operations across fleets of GPU servers and rack-scale systems. This is a software-first role at the intersection of **distributed systems**, **production reliability**, and **hardware-aware automation**. --- ### What You’ll Do - Design, build, and operate **Go-based services** that manage the lifecycle of large-scale GPU data center infrastructure. - Build automation for **data center bring-up**, hardware discovery, health monitoring, remediation, and production operations. - Develop reliable **APIs, services, and workflows** for managing BMCs, firmware state, server health, and rack-level infrastructure. - Improve **observability, alerting, and operational tooling** to detect, understand, and resolve production issues quickly. - Translate incidents and hardware failure modes into **software improvements** that increase platform resilience. - Partner with hardware-adjacent, infrastructure, operations, and software teams to design systems that work safely at fleet scale. - Provide technical leadership via **design reviews, code reviews, architectural guidance, and mentorship**. - Make pragmatic architecture decisions balancing **reliability, simplicity, scalability, and operational burden**. --- ### Who You Are - **B.S./M.S./PhD** in Computer Science (or related) or equivalent experience. - **8+ years** of software engineering experience with strong focus on infrastructure/cloud engineering and distributed systems (especially in large-scale datacenter/cloud environments). - Expert in **Go** and proven experience building **REST/gRPC APIs** for mission-critical platforms. - Strong background architecting and scaling **Kubernetes infrastructure** and distributed services. - Proven success mentoring engineers, leading technical projects, and influencing engineering strategy. - Experience contributing to and collaborating with **open source** communities. - Data-driven approach to reliability, optimization, and continuous improvement. - Excellent communicator with technical and non-technical stakeholders. - Hands-on experience with observability stacks (**Prometheus, Grafana, PromQL**), **CI/CD**, and operating large fleets of GPU servers. - Track record leading incident response, postmortems, and driving service reliability. **Nice to Have** - Working knowledge of **Kafka, ClickHouse, CRDB**. - **DMTF, RedFish APIs**, and GPU servers. --- ### Compensation & Benefits - Base salary range: **$207,000 to $275,000** (based on knowledge, skills, experience, and market location). - Total rewards may include **discretionary bonus, equity awards, and comprehensive benefits** (eligibility-based). --- ### Benefits (US full-time) - Medical, dental, and vision insurance — **100% paid** by CoreWeave - Company-paid Life Insurance - Voluntary supplemental life insurance - Short and long-term disability insurance - Flexibl
Listing freshness
CronJobs last confirmed this listing 11h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.