CronJobs

devops-sre jobs

Staff Software Engineer - Observability

Cerebras · Sunnyvale, CA

hybridstaffPosted Sep 29, 2026GoC++RustJavaPythonOpenTelemetryPrometheusGrafana

Apply on the employer site

About this role

## Staff Software Engineer — Observability ### About the team The Cerebras Inference team’s mission is to deliver the world’s most performant, secure, and reliable enterprise-grade AI service. You’ll help build and operate large-scale distributed systems that power AI inference at unprecedented speed and efficiency. ### What you’ll do - Design and implement observability instrumentation across services and platforms - Build and maintain telemetry pipelines for metrics, logs, and traces at scale - Develop internal observability platforms, libraries, and tooling - Define and operationalize SLIs, SLOs, and alerting strategies - Partner with engineers to make systems debuggable by design - Reduce MTTR by enabling fast root-cause analysis during incidents - Create clear, actionable dashboards and alerts that reflect real system health - Balance telemetry signal vs. cost, noise, and performance impact - Improve the developer experience around observability and debugging ### Qualifications **Core Engineering Skills** - Strong experience in backend or systems software engineering - Proficiency in one or more of: Go, C++, Rust, Java, Python - Solid understanding of distributed systems, networking fundamentals, and concurrency/performance tradeoffs **Observability & Reliability Experience** - Hands-on experience with metrics, logs, and distributed tracing - Production monitoring and alerting experience - Familiarity with tools such as OpenTelemetry, Prometheus, Grafana, Datadog/Elastic/Jaeger/Tempo (or similar) - Experience designing high-signal alerts, scalable telemetry pipelines, and service-level indicators/objectives **Preferred Qualifications** - Experience in high-performance computing, AI/ML systems, or inference platforms - Hardware-aware observability (accelerators, GPUs, custom hardware) - Prior SRE or platform engineering background - Experience debugging large-scale production incidents - Building internal developer platforms or shared libraries ### Why join Cerebras Cerebras builds breakthrough AI hardware and infrastructure—serious software teams building their own hardware. You’ll work on a fast-growing platform with a focus on performance, reliability, and real-world impact. **Links** - OpenAI partnership: https://openai.com/index/cerebras-partnership/ - Learn more / apply: https://www.cerebras.ai/join-us - Privacy/CCPA: https://www.cerebras.net/privacy/

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord