CronJobs

devops-sre jobs

Observability - Site Reliability Engineer

Peraton · Chantilly, VA, US

remoteunknown$112,000–$179,000Posted Oct 11, 2026OpenTelemetryKubernetesPrometheusGrafanaCloudWatchOpenSearchdistributed tracinginfrastructure-as-code

Apply on the employer site

About this role

## Observability / Site Reliability Engineer ### Responsibilities - Design, implement, and operate telemetry and reliability capabilities across enterprise infrastructure, platform services, applications, and workloads. - Implement centralized metrics, logs, traces, dashboards, alerts, and reliability practices using **OpenTelemetry** or equivalent approved standards/tools. - Establish standardized collection patterns for metrics, logs, and distributed traces across enterprise environments. - Integrate cloud, Kubernetes, platform, application, and infrastructure telemetry into centralized observability services. - Maintain telemetry pipelines, retention, routing, and data-quality controls. - Develop and maintain instrumentation and telemetry configurations for effective monitoring and troubleshooting. **Reliability & Performance Engineering** - Build dashboards, alerts, service-health views, and operational metrics for performance and availability visibility. - Define and monitor **Service Level Objectives (SLOs)**, availability indicators, latency, error rates, saturation, and capacity measures. - Analyze performance, capacity, utilization, and operational trends to identify reliability issues and corrective actions. - Apply SRE practices to improve availability, performance, scalability, and operational resilience. - Identify opportunities to improve detection, response, and recovery across enterprise services. **Operations & Automation** - Support incident diagnosis, cross-service troubleshooting, and root-cause analysis for complex observability/reliability issues. - Automate recurring monitoring, alerting, reporting, and remediation activities where appropriate. - Partner with operations and engineering teams to improve incident detection/response and reduce recurring issues. - Participate in Tier 3/4 escalation and on-call support for observability, platform, and reliability issues. - Maintain operational procedures, troubleshooting documentation, and knowledge articles. - Drive continuous improvement of monitoring and reliability practices based on operational data and incident trends. ### Location - **Fully onsite** in **Chantilly, VA** ### Qualifications **Required** - **Active TS/SCI clearance with CI Polygraph** - ~**5–8 years** of experience in observability, SRE, operations, platform engineering, cloud engineering, or related technical discipline - **5 years** with a BS/BA (additional **4 years** may be considered in lieu of a degree) - Hands-on experience with **metrics, logging, tracing, monitoring, and alerting** - Experience with **Kubernetes** and **cloud** environments - Strong troubleshooting, systems analysis, and performance-analysis skills - Experience operating/supporting production services and responding to technical incidents - Experience with **scripting, automation, or infrastructure-as-code** - Strong analytical, problem-solving, and communication skills **Desired** - Experience with one or more of: - **OpenTelemetry** - **Prometheus & Grafana** - **CloudWatch, OpenSearch**, or comparable observability platforms - Distributed tracing / APM - **SLO/SLA** and error-budget practices - Observability for large-scale/distributed enterprise platforms - Automated incident response/remediation - Observability in secure/highly regulated environments ### Target Salary Range - **$112,000 – $179,000** (typical range; final offer depends on factors including experience, education, location, and con

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord