Site Reliability Engineer II
Onapsis · Dallas, Texas, United States
About this role
**About the job** Onapsis eliminates the blind spot in cybersecurity for business-critical applications. The company helps nearly 30% of the Forbes Global 100 understand threats and risks across SAP and Oracle landscapes—on-prem, in the cloud, or hybrid. We’re looking for a **Site Reliability Engineer II** to join our global engineering team. In this role, you’ll apply software engineering principles to operations to ensure our cloud platform and distributed security products are **highly available, scalable, and resilient**. You’ll be accountable for **monitoring, inspecting, troubleshooting, and resolving** service and product issues, while partnering with engineering teams to improve **telemetry** and **operation automations**. Rather than manual operational maintenance, you’ll **write code**, build **automation**, and design **observability frameworks** to eliminate toil and prevent system failures. You’ll work alongside product development teams to embed reliability into the software lifecycle from day one. --- **What you will be doing** - Take shared full-stack ownership of Onapsis production environments, balancing active operational response with modern software engineering practices. - Develop a deep, end-to-end understanding of system architecture, dependencies, and service behaviors to maximize performance, scalability, and resilience. - Split time between live production operations and engineering initiatives that improve long-term stability. - Diagnose complex issues across distributed cloud services and stateful infrastructure during incidents. - Design, develop, and maintain automation tooling to eliminate repetitive toil, optimize monitoring/telemetry, and improve operational efficiency. - Apply a software lens to operational friction—turning vulnerabilities and outages into permanently solved engineering problems. --- **Requirements** - Bachelor’s or Master’s degree in Computer Science (or related) **or equivalent experience**. - **2–4 years** experience in SRE, DevOps, or Cloud Engineering supporting production distributed systems in cloud environments. - Beginner programming skills: **Java and Python**. - Intermediate knowledge of SRE observability practices, including **APM**, **error budget** definition/tracking, and connecting **monitoring, logging, and tracing**. - Understanding of **SLI/SLO/SLA** methodologies. - Beginner knowledge of **Infrastructure as Code** and **Terraform**. - Beginner knowledge of **containers** and orchestration platforms such as **Kubernetes**. - Intermediate knowledge of scripting/automation with **Bash or Python**. - Experience using **AI coding assistants** for daily tasks, debugging, and rapid prototyping. - Experience managing **Linux OS** (Debian/Ubuntu/OpenSuse), **RabbitMQ**, **PostgreSQL**, and AWS infrastructure such as **EC2** or **EKS**. - Intermediate knowledge troubleshooting complex software and networking issues. - Proven ability to quickly learn new technical domains and train others. - Strong verbal and written communication skills. --- **Desired skills or interests in** - Ability to own end-to-end features, service integrations, database schema design, and operational telemetry. - Secure software development best practices. - Knowledge of **TDD**, **CI/CD** tooling, and **Agile** methodologies. - Professional software engineering practices across the full SDLC (coding standards, code reviews, source control, build processes, testing, and operations). - Experienc
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.