Sr. Sustaining and Forward Deployed Engineer
Abacusinsights · United States
About this role
**Sr. Sustaining and Forward Deployed Engineer** **About Abacus Insights** Abacus Insights is transforming how data works for health plans—making healthcare data usable so care and cost decision-makers can act faster, with confidence. We help health plans break down data silos to create a single, trusted data foundation that powers better outcomes, reduces waste, and improves experiences for members and providers. Backed by $100M from top investors, we’re building a platform that enables GenAI use cases by delivering clean, connected, reliable healthcare data. --- ## About the Role The **Senior Site Reliability Engineer – Forward Deployed (AWS & Databricks)** is a senior individual contributor responsible for **production operations, incident response, and post-launch system reliability** across Abacus Insights’ platform. This role blends: - SRE and production operations - SWAT-style deep technical problem solving - Forward-deployed, customer-facing technical work - Hands-on software development and automation You’ll own the most complex, ambiguous, and high-impact production issues—especially those involving **AWS infrastructure, Databricks workloads, and large-scale data pipelines**. You’ll work directly with customers on escalations and deployments, and ensure learnings translate into durable product and platform improvements. --- ## Your Day to Day ### Production Operations & Incident Response - Act as a **senior technical escalation point** during production incidents - Lead **real-time incident triage, mitigation, and recovery** - Drive **root cause analysis (RCA)** focused on systemic, long-term fixes - Identify recurring failure patterns and push for architectural/operational improvements - Partner with **Customer Success and Engineering** to manage customer impact ### Sustaining Engineering & Post-Launch Ownership - Own **post-launch reliability, stability, and operational quality** of core systems - Investigate and resolve complex field issues and production defects - Ensure incident/customer escalation fixes are **upstreamed into the core product** - Improve operational readiness via **runbooks, monitoring, and alerting** - Reduce operational toil by converting manual work into **automation** ### Forward Deployed / Customer-Facing Engineering - Engage directly with strategic customers to solve real-world production challenges - Support complex deployments, integrations, and escalations in customer environments - Serve as a trusted technical partner during high-impact issues - Translate customer learnings into concrete product/platform/operational improvements - Contribute reusable tools, playbooks, and best practices to accelerate future deployments ### AWS & Databricks Technical Expertise - Serve as a subject matter expert for **AWS-hosted production systems** - Troubleshoot across: - AWS compute, storage, networking, IAM, and security - Databricks jobs, clusters, and Spark-based data pipelines - Debug performance degradation, scalability issues, job failures, and data correctness problems - Partner with platform/data teams to harden systems for reliability, scale, and operability ### Software Development & Automation - Write production-quality code to: - Automate operational workflows - Improve reliability and observability - Eliminate manual intervention and reduce incident frequency - Contribute primarily in **Python** (with exposure to JVM-based systems as needed) - Review code with emphasis on **op
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.