CronJobs

devops-sre jobs

Sr. Sustaining and Forward Deployed Engineer

Abacusinsights · United States

remoteseniorPosted Oct 2, 2026PythonAWSDatabricksSparkKubernetesEKSIAMCI/CD

Apply on the employer site

About this role

**Sr. Sustaining and Forward Deployed Engineer** **About Abacus Insights** Abacus Insights is transforming how data works for health plans—making healthcare data usable so care and cost decision-makers can act faster, with confidence. We help health plans break down data silos to create a single, trusted data foundation that powers better outcomes, reduces waste, and improves experiences for members and providers. Backed by $100M from top investors, we’re building a platform that enables GenAI use cases by delivering clean, connected, reliable healthcare data. --- ## About the Role The **Senior Site Reliability Engineer – Forward Deployed (AWS & Databricks)** is a senior individual contributor responsible for **production operations, incident response, and post-launch system reliability** across Abacus Insights’ platform. This role blends: - SRE and production operations - SWAT-style deep technical problem solving - Forward-deployed, customer-facing technical work - Hands-on software development and automation You’ll own the most complex, ambiguous, and high-impact production issues—especially those involving **AWS infrastructure, Databricks workloads, and large-scale data pipelines**. You’ll work directly with customers on escalations and deployments, and ensure learnings translate into durable product and platform improvements. --- ## Your Day to Day ### Production Operations & Incident Response - Act as a **senior technical escalation point** during production incidents - Lead **real-time incident triage, mitigation, and recovery** - Drive **root cause analysis (RCA)** focused on systemic, long-term fixes - Identify recurring failure patterns and push for architectural/operational improvements - Partner with **Customer Success and Engineering** to manage customer impact ### Sustaining Engineering & Post-Launch Ownership - Own **post-launch reliability, stability, and operational quality** of core systems - Investigate and resolve complex field issues and production defects - Ensure incident/customer escalation fixes are **upstreamed into the core product** - Improve operational readiness via **runbooks, monitoring, and alerting** - Reduce operational toil by converting manual work into **automation** ### Forward Deployed / Customer-Facing Engineering - Engage directly with strategic customers to solve real-world production challenges - Support complex deployments, integrations, and escalations in customer environments - Serve as a trusted technical partner during high-impact issues - Translate customer learnings into concrete product/platform/operational improvements - Contribute reusable tools, playbooks, and best practices to accelerate future deployments ### AWS & Databricks Technical Expertise - Serve as a subject matter expert for **AWS-hosted production systems** - Troubleshoot across: - AWS compute, storage, networking, IAM, and security - Databricks jobs, clusters, and Spark-based data pipelines - Debug performance degradation, scalability issues, job failures, and data correctness problems - Partner with platform/data teams to harden systems for reliability, scale, and operability ### Software Development & Automation - Write production-quality code to: - Automate operational workflows - Improve reliability and observability - Eliminate manual intervention and reduce incident frequency - Contribute primarily in **Python** (with exposure to JVM-based systems as needed) - Review code with emphasis on **op

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord