Site Reliability Engineer III (DBA)
Backblaze · Remote - US
About this role
## About Backblaze Backblaze is the object storage leader in the open cloud movement, helping customers unlock budgets, unburden administrators, and unleash innovators with cloud storage built to break free from restrictive, overpriced legacy solutions. Founded in 2007, Backblaze scaled with less than $3M in outside funding until 2021, when it completed a traditional IPO on the Nasdaq. Today, Backblaze generates over $100M in revenue and manages 3B+ GB of data storage for 500K+ customers across 175+ countries. ## About the Role We’re seeking a **Site Reliability Engineer III (DBA)** to ensure the stability, scalability, and reliability of our production database systems—primarily **Vitess (distributed MySQL)** and **Cassandra**—along with the rest of our production services and infrastructure. This role includes the same **on-call, incident response, and service ownership** expectations as other SRE IIIs, with database systems as the area of deepest technical ownership. Because our SRE Database Engineering function is new, you’ll also help establish its operational foundation by designing database architecture and creating **runbooks, escalation guidance, procedures, and training materials** for Level 1 and Level 2 SRE Database Engineers. ## What You’ll Do ### Database Architecture & Administration - Design, deploy, and own highly available database architecture for **Vitess** and **Cassandra** - Establish and document operational procedures, runbooks, and escalation guidance for Level 1 and Level 2 SRE Database Engineers - Optimize database performance via query tuning, indexing strategies, schema design, and capacity planning - Own backup, recovery, replication, and disaster recovery strategies - Perform and validate disaster recovery testing and database recovery procedures - Drive database security, access control, patching, hardening, and compliance practices - Partner with DBA and Data Infrastructure teams on resharding, capacity planning, replication, and architecture decisions for sharded MySQL environments ### Service Reliability & Operations - Support availability and durability of critical services across production environments - Monitor service health using **SLIs, SLOs, error budgets**, and monitoring/logging/alerting platforms - Partner with service owners to define and improve SLIs, SLOs, error budget policies, and alerting - Participate in on-call rotations, incident response, root cause analysis, and post-incident reviews - Serve as an escalation point for complex database production incidents - Follow ITIL/OSS processes (incident, change, problem, and capacity management) - Take ownership of operational issues and drive projects from discovery through resolution ### Automation & Tooling - Develop automation to reduce manual intervention and operational toil - Contribute to monitoring/logging/alerting frameworks including **Prometheus, Grafana, Catchpoint, and ELK** - Integrate operational runbooks and incident response workflows with **FireHydrant** - Work with CI/CD, configuration management, and infrastructure as code tools including **Terraform, Ansible, and Jenkins** - Develop scripts using **Bash, Python, Go**, or similar technologies - Operate and troubleshoot containerized production environments using **Kubernetes and Docker** - Work within Kubernetes and Vitess environments using tools such as **kubectl, mysqlsh, and Vitess keyspaces** ### Project Management - Lead Production Readiness Reviews (PRRs)
Listing freshness
CronJobs last confirmed this listing 57m ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.