CronJobs

devops-sre jobs

Site Reliability Engineer, Infrastructure Platforms — UK (Intermediate to Senior Staff)

GitLab · Remote

remoteseniorPosted Sep 8, 2026KubernetesTerraformGoAWSGCPCI/CDInfrastructure as Code

Apply on the employer site

About this role

**Site Reliability Engineer, Infrastructure Platforms — UK** *Intermediate to Senior Staff* **About the Role** Site Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale. You'll combine software engineering with operational excellence, applying sound engineering principles, automation, and continuous improvement to build, operate, and evolve production infrastructure. This is a single application for SRE opportunities across Infrastructure Platforms teams. We evaluate skills holistically and match you to the opportunity that best aligns with your experience and our hiring needs. **What You'll Do** • Keep user-facing services and production systems reliable, scalable, and efficient • Build automation and tooling that reduces toil and replaces manual work • Operate and troubleshoot production systems on Kubernetes • Write and maintain infrastructure as code, shipping changes safely through CI/CD and GitOps • Participate in on-call, triage alerts, and improve runbooks • Contribute to the observability stack using metrics, logs, and SLOs • Take part in incident response and post-incident reviews • Document runbooks, architecture decisions, and turn learnings into repeatable practices **What You'll Bring** • Experience keeping production systems reliable with an operations mindset and real software engineering practice • Experience building net-new infrastructure tooling and automation (e.g., Terraform modules, Kubernetes operators) • Ability to read, debug, and reason about code (Go or Ruby experience helpful) • Experience with infrastructure as code and Kubernetes at an appropriate depth • Hands-on experience with GCP or AWS • Familiarity with observability practices: metrics, logging, alerting, SLOs/SLIs • Comfort with on-call and incident response with structured troubleshooting • Strong written communication and ability to work async in a distributed environment • Track record of using automation and AI to reduce toil • Alignment with GitLab's values **About Infrastructure Platforms** Responsible for the availability, reliability, performance, and scalability of GitLab's user-facing services. Spans teams across Production Engineering, GitLab Dedicated, GitLab Delivery, and Developer Experience. Globally distributed, remote-first, working asynchronously with a focus on automation over toil. **Note:** This position is open to candidates based in the United Kingdom only.

Listing freshness

CronJobs last confirmed this listing 23h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord