CronJobs

backend jobs

Senior Software Engineer - Incident Insights & Readiness

Datadog · Boston, Massachusetts, USA; New York, New York, USA

hybridsenior$192,000–$192,000Posted Oct 1, 2026GoPythonTypeScriptKubernetes

Apply on the employer site

About this role

**Senior Software Engineer - Incident Insights & Readiness (Datadog)** We’re on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams—at high scale, with always-on alerting, metrics visualization, logs, and application tracing. The **Incident Insights & Readiness SRE** team fosters a resilient culture by using incidents as learning opportunities and catalysts for growth. You’ll help build the software, tooling, and operational frameworks that prepare teams to respond to incidents, learn from them, and continuously improve reliability. **What you’ll do** - Own and improve the on-call experience by establishing best practices and building platforms to support on-call rotations and compensation. - Define how we respond to incidents; lead design and implementation of software to streamline the process; collaborate with product teams to improve incident response across Datadog. - Contribute to the post-mortem process: help write post-mortems, identify opportunities to reduce friction, and enhance learning value (including a weekly postmortem reading group). - Support incident reviews that emphasize learning and blamelessness; help share learnings across the organization to improve resilience. - Provide technical leadership and day-to-day coaching through design reviews, collaborative problem-solving, and operational excellence best practices. - Train on-callers in incident and post-mortem processes—onboard newcomers and refresh knowledge for existing engineers. - Lead cross-functional initiatives by embedding with engineering teams to understand challenges and drive lasting improvements. **Who you are** - At least **5 years** building software that solves real user problems; experience designing features and collaborating on code/technical design reviews. (Primarily **Go** and **Python**, with some **TypeScript**.) - Experience building/operating **distributed systems**; familiarity with **Kubernetes** and complex failure modes. - Ability to independently own ambiguous technical problems from design through delivery, balancing long-term quality with pragmatic execution. - Experience analyzing incidents, identifying systemic risks, and driving engineering improvements from operational learnings. - Experience participating in on-call rotations and improving incident response processes; experience as an incident commander/coordinator is a plus. - Empathy, collaboration, and strong communication skills in English. - Experience mentoring engineers, driving cross-functional initiatives, and influencing technical direction without relying on organizational authority. - Open to candidates from varied backgrounds (software engineering, SRE, production engineering, infrastructure, and other roles focused on reliability and incident response). **Benefits and Growth** - New hire stock equity (RSUs) and employee stock purchase plan (ESPP) - Continuous professional development, product training, and career pathing - Intradepartmental mentor and buddy program - Inclusive culture and access to Community Guilds (employee resource groups) - Inclusion Talks (internal panel discussions) - Free, global mental health benefits for employees and dependents age 6+ - Competitive global benefits *Benefits may vary by country and employment type.* **Compensation (estimated)** - **$192,000 — $240,000 USD** (yearly) **About Datadog** Datadog is the leading observability and security platform

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord