Senior Software Engineer II, Observability
CoreWeave · New York, NY
About this role
**CoreWeave — Senior Software Engineer II (Observability)** **About the Role** CoreWeave is seeking **Senior Software Engineers** specializing across the pillars of **Observability**. You’ll help build the best-performing systems to market—whether your focus is **metrics, logging & tracing, or telemetry pipelines & visualization**—so CoreWeave and its customers can **understand, troubleshoot, and optimize complex AI systems**. **What You’ll Do** - **Design, build, and own core observability infrastructure**, including scalable and reliable logging, metrics, and tracing platforms. - **Develop and operate high-throughput telemetry pipelines** to ingest, transform, and expose observability data—ensuring reliability, security, and transparent data migrations. - **Tackle observability challenges at extreme scale**, supporting clusters of thousands of GPUs, petabyte-scale telemetry, and high-cardinality workloads. - **Continuously improve performance, security, reliability, and scalability** through software enhancements, automation, and new features. - Join the team’s **on-call rotation** for critical production systems, focusing on root cause analysis and durable fixes. - **Collaborate with internal engineering teams** using a platform-as-a-product mindset to embed observability best practices and tooling. - Contribute to the **overall observability strategy**, influencing platform direction and customer experience. **Who You Are** - **5+ years** of experience in software or infrastructure engineering, with a track record building and operating large-scale distributed systems. - Proficient in **Go (primary)** or **Python**, writing clean, resilient, testable production code. - Hands-on **Kubernetes** experience in production (containerization, microservices) and familiarity with observability challenges. - Demonstrated ability to **design, build, and deliver robust, scalable systems** with operational excellence and strong testing practices. - Ability to **analyze and decompose complex problems** in elastic, distributed architectures. - Comfortable with **Helm and YAML-based configuration** (templating, automation, infrastructure-as-code). - **Customer-obsessed, platform-minded** approach—providing infrastructure as a service and applying a product lens to platform scale problems. - Experience participating in an **on-call rotation** for critical production systems. **Preferred Qualifications** - Direct experience designing/operating/scaling observability platforms (e.g., **Loki, ClickHouse, Elasticsearch, Prometheus, VictoriaMetrics, Grafana, Thanos**). - Familiarity with **data streaming systems** (e.g., **Kafka, Kafka Connect**) for observability pipelines. - Experience **automating and provisioning infrastructure** (e.g., **Terraform**). - Experience with **OpenTelemetry** for unified telemetry collection and instrumentation. - Experience with **modern AI platforms/workloads** (large-scale training/inference, GPU infrastructure, MLOps) is a plus. **Compensation & Benefits** - **Base salary range:** **$182,000 to $242,000** (final offer based on knowledge, skills, experience, and market/location). - Total rewards may include **discretionary bonus, equity awards, and benefits** (eligibility-based). - US full-time benefits may include: **100% paid medical/dental/vision**, life insurance, disability coverage, **FSA/HSA**, tuition reimbursement, **ESPP**, Spring Health mental wellness benefits, Carrot family-forming support, paid parent
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.