Senior AI/ML Operations Engineer
Abacusinsights · United States
About this role
**Senior AI/ML Operations Engineer** **About Us** Abacus Insights is transforming how data works for health plans. Our mission is simple: make healthcare data usable—so the people responsible for care and cost decisions can act faster, with confidence. We help health plans break down data silos to create a single, trusted data foundation. That foundation powers better decisions—so plans can improve outcomes, reduce waste, and deliver better experiences for members and providers alike. Backed by $100M from top investors, we’re tackling big challenges in an industry that’s ready for change. **About the Role** The Senior AI/ML Ops Engineer is a senior individual contributor responsible for the infrastructure, pipelines, and operational reliability that power both classical machine learning (ML) and generative AI (GenAI)/agentic systems on Databricks and Snowflake. You will own the reliability of ML and GenAI systems in production—from ML/AI-specific pipeline and deployment workflows through model lifecycle management to the infrastructure behind retrieval and agentic tooling—taking AI engineering and data science work from prototype to reliable production systems. **Your day to day** **Platform & Pipeline Engineering** - Deploy and promote ML and GenAI models, pipelines, and code across environments using CI/CD infrastructure maintained by Infra/DevOps - Develop and promote reusable deployment patterns and tooling to reduce effort to stand up new AI/ML use cases and client-specific deployments - Build and maintain data pipelines supporting both classical ML and GenAI workloads (ingestion → feature engineering → serving) - Operate within Databricks and Snowflake governance frameworks (e.g., Unity Catalog access controls, environment boundaries) to ensure secure, compliant promotion of code, data, and models - Independently diagnose and resolve production issues across pipelines, infrastructure, and model-serving systems **Classical ML Operations** - Automate and monitor production ML inference and feature engineering workflows, including alerting and incident response - Own model lifecycle management using a model registry tool such as MLflow, along with Unity Catalog (experiment tracking, model registration, versioning, and controlled promotion) **GenAI & Agentic Infrastructure** - Build and maintain infrastructure for retrieval-augmented generation (RAG) systems, including vector search indexing and retrieval pipelines - Deploy, host, and maintain MCP servers and tool integrations for agentic applications - Build and maintain evaluation infrastructure for AI systems and contribute to evaluation methodology - Support agent observability: logging, tracing, and monitoring for agent and model behavior in production **Cross-Functional** - Partner with business stakeholders to scope data and feature requirements - Coordinate with Software Engineering, Data Engineering, Data Science, Security, and DevOps on infrastructure changes and shared platform needs - Mentor junior engineers on platform practices and operational standards - Occasionally contribute to customer-specific implementation work as part of a broader team **What you bring to the team** - 5+ years of experience in AI/ML engineering, MLOps, or a closely related discipline - Hands-on depth in at least one of: - **AI/GenAI:** agentic frameworks (e.g., LangChain), RAG systems, vector search, MCP/tool-integration protocols, model serving/gateway layers, evaluation design -
Listing freshness
CronJobs last confirmed this listing 1d ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.