Senior Software Engineer I (Observability Platform)
Smartsheet · -REMOTE, USA-
About this role
**Senior Software Engineer I (Observability Platform)** Smartsheet’s Observability Engineering team builds the platform that helps the company “see itself”—collecting, modeling, storing, and analyzing metrics, logs, distributed traces, and events across a global, multi-region infrastructure. As a Senior Software Engineer I, you’ll build the observability platform end-to-end (instrumentation libraries, telemetry pipelines, SLO/alerting systems, and dashboards as code) and connect it to the systems where engineering work happens—CI/CD, service catalog, incident management, ticketing, chat, and automated remediation. --- ### You Will - **Architect and build the observability platform:** Ship end-to-end capabilities for metrics, logs, distributed traces, and events across regions/environments. - **Run observability as an internal product:** Provide published interfaces, versioned client libraries, clear deprecations, and SLOs for the platform itself. - **Drive instrumentation with OpenTelemetry:** Build shared instrumentation libraries, collector deployments, semantic conventions, context propagation, and sampling strategies. - **Engineer telemetry pipelines at scale:** Build reliable, high-volume collection/enrichment/redaction/routing with backpressure and tiered retention. - **Connect the platform end to end:** Integrate with CI/CD, service catalog, incident management, ticketing, chat, and feature-flag systems via REST APIs, webhooks, and event-driven services. - **Build automated and self-healing remediation:** Create event-driven/agent-assisted workflows with human-in-the-loop approval for consequential actions. - **Make reliability measurable:** Implement SLOs, error budgets, and golden-signal alerting as code to reduce alert noise. - **Build the golden path:** Own the paved road for new service instrumentation (templates, versioned Terraform modules, CI checks). - **Instrument the platform itself:** Track coverage, onboarding time, time to first useful dashboard, query performance, and cost per service. - **Build the enablement layer:** Documentation-as-code, reference architectures, worked examples, workshops/office hours/game days. - **Raise the technical bar:** Lead reviews and architecture discussions; mentor on signal design and cost-aware instrumentation. - **Apply AI where it earns its place:** Use AI to improve engineering efficiency and help make Smartsheet’s AI/agentic systems observable. - **Turn incidents into durable improvements:** Participate in on-call, drive RCA, and close incidents with instrumentation/automation/lessons learned. --- ### You Have - **5+ years** building and operating highly scalable, highly available distributed systems, platform services, or observability infrastructure. - **5+ years** programming in **Go, Python, Java/Kotlin, or TypeScript/Node.js**, with comfort moving between backend services and operational tooling. - Experience building **internal platforms / developer tooling** with a **product mindset**. - Hands-on depth across **metrics, logs, and distributed tracing** on a major commercial or open-source observability platform, including cardinality/sampling/cost mechanics and SLO/alerting. - Practical **OpenTelemetry** experience (collectors, auto/manual instrumentation, semantic conventions, context propagation). - Experience building **high-volume telemetry pipelines** (shippers, streaming transport, transformation/enrichment, and search/time-series backends). - Strong **REST API
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.