CronJobs

devops-sre jobs

Senior Site Reliability Engineer

Alembic · Atlanta Perimeter

onsitesenior$200,000–$225,000Posted Sep 16, 2026PythonKubernetesTerraformAWSDockerAnsibleBashPrometheus

Apply on the employer site

About this role

**Senior Site Reliability Engineer** **About the Role** We’re looking for an experienced Site Reliability Engineer (SRE) to help scale our platform with reliability, observability, and operational excellence. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure powering our core platform—including data pipelines, ML workloads, and real-time analytics systems. This is a hands-on, high-impact role with visibility across the stack. **Location**: Onsite in **Dunwoody, GA** --- **Key Responsibilities** - Design, build, and maintain scalable infrastructure for real-time analytics and machine learning workloads - Improve reliability and performance through automation, observability, and proactive capacity planning - Own and evolve CI/CD pipelines, deployment automation, rollback mechanisms, and configuration management - Implement and maintain monitoring, alerting, and incident response processes (SLOs, runbooks, on-call rotations) - Collaborate across engineering and data science teams to drive a culture of performance and reliability - Ensure security, compliance, and operational readiness across cloud infrastructure - Lead post-incident analysis and continuous improvement initiatives --- **What Will Help You Succeed** - 8+ years of experience in SRE, DevOps, or infrastructure engineering - 5+ years of experience with datacenter operations and/or system and network administration - Experience with containerization (**Docker**) and orchestration (**Kubernetes**) - Strong Linux systems, networking, and performance tuning skills (including strong command-line/SSH/shell proficiency) - Solid understanding of infrastructure-as-code (e.g., **Terraform**, **Ansible**) - Programming skills for IaC and scripting (e.g., **Terraform**, **Ansible**, **Bash**, and/or **Python**) - Experience with monitoring/observability stacks (e.g., **Prometheus**, **Grafana**, **Datadog**, **ELK**, **OpenTelemetry**) - Proficiency with CI/CD tools and pipelines (e.g., **GitHub Actions**, **ArgoCD**) - Ability to debug complex systems and automate solutions via scripting - Excellent communication skills and ability to work cross-functionally --- **Nice-to-Have** - Experience with cloud and managed services (e.g., **AWS**) - Experience supporting data-intensive platforms (e.g., **Spark**, **Airflow**, **Kafka**) - Familiarity with security practices for cloud-native applications and infrastructure - Experience in high-compliance or **SOC-2** environments --- **What You’ll Get** - Ownership of mission-critical infrastructure in a company solving real-world enterprise problems - A front-row seat to a high-performance engineering culture - The ability to influence how the platform scales—from deployment to incident management - An environment that values curiosity, accountability, and impact

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord