CronJobs

devops-sre jobs

Engineering Manager, Site Reliability

Instacart · United States - Remote

remotesenior$214,000–$214,000Posted Oct 1, 2026

Apply on the employer site

About this role

**Engineering Manager, Site Reliability** **About Instacart** Instacart is transforming the grocery industry by building technology that connects customers, shoppers, retailers, and brands through a reliable, convenient online marketplace. The Site Reliability Engineering team helps ensure this experience remains resilient, scalable, and dependable as Instacart grows. **Flex First** Instacart is a Flex First organization—choose where you do your best work (home, office, or your favorite coffee shop) while staying connected through regular in-person events. Learn more: https://www.instacart.careers/flex-first --- ## **About the Job** We’re seeking an Engineering Manager to lead a team of Site Reliability Engineers responsible for the systems, tools, and practices that support the reliability of Instacart’s technology platform. You’ll manage and develop engineers while partnering across the company to improve availability, scalability, performance, observability, and operational excellence. You’ll help establish sound engineering practices, guide complex technical initiatives, and create an environment where teams can build and operate dependable systems. ### **What you’ll do** - Lead, mentor, and develop a team of Site Reliability Engineers; set clear goals and support career growth. - Set technical direction and priorities to improve reliability, scalability, availability, performance, and operational readiness. - Partner with engineering, product, security, infrastructure, and other teams to define reliability standards and influence system design. - Drive incident management and operational excellence (incident response, post-incident learning, SLOs, capacity planning, observability, and continuous risk reduction). - Promote automation and self-service tooling to reduce operational toil and improve deployment confidence. - Balance near-term operational needs with long-term investments in a fast-changing, high-growth environment. - Communicate clearly with technical and non-technical stakeholders about reliability risks, status, tradeoffs, and decisions. This role requires comfort with ambiguity, calm decision-making during incidents, and a blame-free approach to learning from failures. --- ## **About You** ### **Minimum Qualifications** - Bachelor’s degree in Computer Science, Engineering, or related field, or equivalent practical experience. - 7+ years of experience in software engineering, infrastructure engineering, Site Reliability Engineering, or related. - 2+ years of experience managing, mentoring, or leading engineering teams. - Professional experience with cloud infrastructure, distributed systems, networking, containers, orchestration platforms, or related production technologies. - Experience leading/participating in production incident response, post-incident reviews, reliability improvements, and operational readiness. - Experience communicating technical risks, priorities, and tradeoffs to engineering leaders and cross-functional stakeholders. ### **Preferred Qualifications** - Experience leading SRE, platform engineering, infrastructure engineering, or developer productivity teams. - Experience operating highly available services at significant scale and improving SLOs, observability, capacity, or disaster recovery. - Experience with infrastructure as code, continuous delivery, monitoring/logging/tracing, and automated remediation. - Experience building reliability programs across multiple teams (shared standa

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord