Senior Site Reliability Engineering
Duolingo · Pittsburgh, PA
About this role
**Our mission** Duolingo’s mission is to develop the best education in the world and make it universally available. **About the role** As a **Senior Site Reliability Engineer**, you’ll work closely with product and platform engineering teams to ensure Duolingo’s distributed systems and products are built, maintained, and operated with extraordinary quality—measurably and at scale. **🧠 You will…** - Collaborate with internal teams to identify sources of instability in distributed systems and drive operational excellence - Support core infrastructure (understand, diagnose, and debug these systems in production) - Provide system design consulting, develop software platforms/frameworks, and conduct launch reviews and root cause analysis - Maintain and document sustainable postmortem/incident response practices - Advocate for and implement changes that improve reliability, scalability, and velocity - Reduce the burden of toil with iterative development of tooling and automation - Collaborate with engineering teams to release new features and become an authority on our services **✅ You have…** - 5+ years of experience in site reliability engineering/DevOps for a product with millions of users - Experience identifying and solving issues in large-scale distributed systems - Experience with **Java, Kotlin, Python, or Go** - Understanding of containerization and orchestration technologies (e.g., **Docker, Mesos, Kubernetes, Nomad**) **⭐ Exceptional candidates will have…** - Experience improving automation and tooling to reduce service maintenance toil - Proven experience driving improvements to incident response processes - Experience assessing reliability and troubleshooting issues in **Dynamo, MySQL, and/or PostgreSQL**
Listing freshness
CronJobs last confirmed this listing 6d ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.