CronJobs

devops-sre jobs

Engineering Manager, Site Reliability Engineering

Replit · Foster City, CA

hybridunknown$250,000–$325,000Posted Sep 25, 2026KubernetesGCPOpenTelemetryGitOpsArgoCDHarnessKargo

Apply on the employer site

About this role

## About Replit Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation. ## About the Role Replit enables people to build software with AI. The systems underneath that experience must support safe production changes, measurable reliability, and predictable performance as usage grows. This Engineering Manager will lead SRE across: - Observability - Incident management - Load testing - Performance engineering - Cloud cost and capacity - Rollout infrastructure You’ll lead and grow an existing team that builds and operates production platforms, working hands-on across application and infrastructure boundaries. This is a software-building leadership role—not simply an incident-management function. ## What You’ll Do - **Observability:** Build and operate metrics, logs, traces, and alerting. Help teams establish meaningful SLOs and use production telemetry to diagnose problems and verify improvements. - **Incident Management:** Own incident tooling and practices, coordinate cross-team response, and turn incident reviews into engineering improvements that reduce recovery time and repeat failures. - **Load Testing:** Build and maintain load/failure testing capabilities. Validate critical paths under expected demand, quantify headroom, and test recovery and production readiness with service owners. - **Performance Engineering:** Lead deep engagements with internal teams on SLOs and end-to-end performance. Use profiling, telemetry, and load tests to identify bottlenecks and deliver improvements with service owners. - **Stay Technically Engaged:** Review designs and production changes, debug difficult failure modes, and use AI coding tools (including Replit) to prototype and automate. Apply rigorous review and verification to AI-generated changes. - **Build and Grow a High-Ownership Team:** Coach engineers, develop technical leaders, manage performance, and hire against agreed needs. Make distributed collaboration, mentoring, and backup coverage deliberate. - **Measure Outcomes and Close the Loop:** Track rollout safety, recovery time, repeat incidents, critical-path latency/throughput, test coverage, and improvements from cost/capacity analysis. Agree success measures and continuing ownership with partner teams. ## What You’ll Bring - **Demonstrated engineering management:** Led and developed engineers, made prioritization and performance decisions, hired thoughtfully, and delivered through a team. - **Production systems depth:** Built and operated distributed systems or reliability platforms; can reason across deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms. - **Safe-change and performance judgment:** Led consequential migrations or incidents and used measurement to diagnose reliability/performance problems; can distinguish symptoms from causes and validate fixes under realistic conditions. - **Platform-product and cross-team judgment:** Can build capabilities other teams adopt, lead hands-on engagements without absorbing every service’s operations, and make clear tradeoffs among reliability, performance, engineering effort, and cost. ## Nice to Have - Experience with GitOps or progressive-delivery platforms (Harness, ArgoCD, Kargo) - Experience with observability, profiling, load-testing, and failure-testin

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord