CronJobs

devops-sre jobs

Senior Site Reliability Engineer

Formationbio · New York, NY; Boston, MA; San Francisco, CA

hybridsenior$185,500–$232,000Posted Sep 18, 2026AWSKubernetesTerraformOpenTofuDockerPythonSnowflakeGitHub

Apply on the employer site

About this role

**Senior Site Reliability Engineer — Formation Bio** ## About Formation Bio Formation Bio is a tech and AI-driven pharma company differentiated by radically more efficient drug development. Advancements in AI and drug discovery are creating more candidate drugs than the industry can progress due to the high cost and time of clinical trials. Formation Bio (founded in 2016 as TrialSpark Inc.) builds technology platforms, processes, and capabilities to accelerate all aspects of drug development and clinical trials. ## About the Position As a **Senior Site Reliability Engineer**, you will build and operate the infrastructure, delivery systems, and operational practices that allow the engineering organization to ship reliable software quickly and safely. You’ll work across cloud infrastructure, developer platforms, observability, and production workloads—including product applications, internal tools, data systems, and ML/AI workloads—taking problems from initial diagnosis through implementation and production adoption. Formation Bio is an **AI-native engineering organization**. You’ll use modern AI tools (including agentic coding systems) as part of daily engineering practice, with the judgment needed to validate output and operate reliable production systems. ## Responsibilities - Own the infrastructure and operational platform for shared engineering workloads (compute, runtime environments, orchestration, deployment, observability, access controls, and reliability). - Build and operate secure, observable, reliable infrastructure for product applications, containerized services, internal tools, data systems, ML pipelines, inference, and agentic software. - Research, develop, and maintain core AWS infrastructure and additional cloud outposts across development, staging, and production (compute, networking, databases, load balancers, secrets management). - Create, review, maintain, and optimize infrastructure as code, CI/CD pipelines, and reusable platform patterns. - Establish strong operational practices: SLOs, monitoring, alerting, runbooks, incident response, advanced diagnostics, root cause analysis, and post-incident follow-through. - Partner with Product Engineering, Data Engineering, and Data Science to evaluate, negotiate, and implement architecture and infrastructure for product software, data systems, model training, and inference. - Use AI tools to accelerate infrastructure development, investigate incidents, improve documentation, build automation, and make operational improvements (while validating output). - Participate in support rotation and incident response; maintain a bias for automation and know when ClickOps is appropriate. - Write and review requirements, design documents, and operating procedures; share knowledge and mentor engineers on infrastructure and SRE fundamentals. ## About You - 5+ years of relevant experience in Site Reliability Engineering, infrastructure, systems, DevOps, or similar. - Production experience operating cloud infrastructure and distributed systems with strong operational and reliability judgment. - Experience with advanced diagnostics, incident response, root cause analysis, observability, and automation. - AWS and Snowflake experience; Azure, GCP, and/or Vercel is a plus. - Working experience with Docker, GitHub, Kubernetes, Python, Terraform or OpenTofu, and virtual networking (Terragrunt is a plus). - Experience supporting production ML/AI workloads, MLOps infrastructure, workflow o

Listing freshness

CronJobs last confirmed this listing 58m ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord