Senior Site Reliability Engineer
Camunda · Remote
About this role
**About Camunda** Camunda is the enterprise platform for agentic orchestration—helping organizations coordinate AI agents, people, and systems across complex, end-to-end business processes. With built-in governance, auditability, and human oversight, Camunda helps move AI from pilots to production safely and at scale. Fully remote and global, Camunda is transforming into an AI-first organization built on its own platform. --- **About the role (Senior Site Reliability Engineer)** You’ll build and maintain reliable, scalable infrastructure that helps developers ship better software faster. You’ll design and operate a Kubernetes-based multi-cloud platform, improve monitoring and observability, and collaborate with product and engineering teams to solve real problems in real time. This is a “you build it, you run it” role—owning the systems that power Camunda, driving automation, and mentoring others. --- **What you’ll be doing** - **Design and maintain infrastructure** - Evolve Camunda’s Kubernetes-based, multi-cloud platform architecture to ensure availability, scalability, and fault tolerance. - Establish configuration best practices and network services teams rely on. - **Build observability that matters** - Implement and improve monitoring and alerting so SREs and developers have clear visibility into system health and performance. - Make it easy for teams to understand what’s happening across the stack. - **Own your systems end-to-end** - Participate in on-call rotations and act as the go-to person for quick fixes. - Create runbooks and automation to turn complex problems into manageable processes. - **Ship improvements with product teams** - Work cross-functionally with product engineering, product management, and support to deliver impactful features. - **Push automation to the next level** - Identify repetitive work and automate it away. - Share learnings so the whole team improves. - **Be the expert others learn from** - Help less experienced engineers tackle complex infrastructure challenges. - Break down problems into clear steps and support skill growth. --- **What you bring** **Must-haves** - **Deep hands-on Kubernetes experience** - Built, deployed, and maintained Kubernetes clusters in production. - Strong understanding of workloads, networking, and storage at scale. - **Infrastructure as Code (IaC)** - Proficient with Terraform (or similar), including versioning, testing, and safe deployments. - **Monitoring & observability experience** - Worked with Prometheus, Grafana, or similar; instrument systems and alert on what matters. - **Incident response & 3rd-level support skills** - Diagnosed complex production issues, communicated clearly under pressure, and performed root-cause analysis. - **Passion for automation and raising the quality bar** - Focus on reliability, maintainability, and clarity. - **Responsible use of AI tools** - Use AI for research, code review, documentation, test generation, and automation. - Validate AI outputs against requirements and maintain human accountability. **Nice-to-haves** - Experience with major cloud providers (e.g., AWS EKS, GCP GKE) - ArgoCD / GitOps workflows - Proficiency in Python, Go, or similar - Experience with SLOs and alerting frameworks --- **Location / work model** - Fully remote (role includes on-call responsibilities) --- **Compensation & benefits (high level)** - Competitive, location-based salary ranges (US/UK/Singapore/Can
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.