Site Reliability Engineer (SRE)
Airapps · San Francisco
About this role
## About Air Apps At Air Apps, we believe in thinking bigger—and moving faster. We’re a family-founded company on a mission to create the world’s first AI-powered Personal & Entrepreneurial Resource Planner (PRP). Born in Lisbon, Portugal (2018), we now have offices in both Lisbon and San Francisco, and we’ve remained self-funded while reaching over 100 million downloads worldwide. ## The Role — Site Reliability Engineer (SRE) As a Site Reliability Engineer (SRE) at Air Apps, you’ll ensure the reliability, availability, and scalability of our systems. You’ll work at the intersection of software development and operations—implementing automation, monitoring, and performance optimization to minimize downtime and improve system resilience. ## Responsibilities - Design and implement scalable, reliable, and fault-tolerant systems across cloud environments. - Develop and maintain observability tools (monitoring, logging, alerting) such as Prometheus, Grafana, Datadog, and ELK. - Automate infrastructure provisioning, deployment, and incident response using Infrastructure as Code (IaC) tools like Terraform or CloudFormation. - Optimize system performance, scalability, and incident response workflows to improve uptime. - Partner with development and DevOps teams to improve system design for reliability. - Perform root cause analysis (RCA) and implement preventative measures to minimize failures. - Ensure high availability via load balancing, failover, and disaster recovery strategies. - Improve CI/CD pipelines to increase deployment speed while maintaining stability. - Optimize cloud cost and resource utilization for AWS, Azure, or GCP. - Participate in on-call rotations to address system failures quickly and minimize downtime. ## Requirements - 4+ years of experience in SRE, DevOps, or System Engineering. - Strong knowledge of cloud platforms (AWS, Azure, or GCP) and cloud-native architectures. - Experience with observability/monitoring tools (Prometheus, Grafana, ELK, Datadog, New Relic). - Proficiency with Infrastructure as Code (Terraform, CloudFormation, or Pulumi). - Hands-on experience with containerization and orchestration (Docker, Kubernetes, Helm). - Strong Linux system administration and networking fundamentals. - Experience with incident management, debugging, and root cause analysis. - Scripting proficiency (Bash, Python, or Go) for automation and monitoring. - Knowledge of load balancing, failover strategies, and distributed systems. - Understanding of security best practices, access control, and compliance requirements. - Strong communication skills and ability to collaborate cross-functionally. ## What Benefits Are We Offering? - Apple hardware ecosystem for work - Annual bonus - Medical insurance (including vision & dental) - Disability insurance (short and long-term) - 401k up to 4% contribution - Air Conference (meet the team, collaborate, and grow) - Transportation budget - Free meals at the hub - Gym membership ## Diversity & Inclusion Air Apps is committed to fostering a diverse, inclusive, and equitable workplace. We welcome applicants from all backgrounds, experiences, and perspectives. ## Application Disclaimer Applicants must submit their own work without any AI-generated assistance. Any use of AI in application materials, assessments, or interviews will result in disqualification.
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.