CronJobs

backend jobs

Applied AI Engineer

Monte Carlo Data · Remote, Americas

remoteunknown$180,000–$240,000Posted Aug 27, 2026PythonLLMsRAGClaude

Apply on the employer site

About this role

## Applied AI Engineer ### About Monte Carlo Monte Carlo is the agent trust platform that unifies data and agent observability to monitor, troubleshoot, and improve production AI systems. As enterprises deploy thousands of agents across business-critical use cases, Monte Carlo provides the reliability infrastructure to support them—from human-guided agents to fully autonomous operations. ### The Role We’re building products that help enterprises determine whether their AI agents can be trusted. You’ll work end-to-end—from an ambiguous problem statement through research, prototyping, proving what works, and integrating it into the platform alongside engineering and data science teams. ### What You’ll Do - **Own an open problem end-to-end**: research and prototype through production; kill what doesn’t work before it becomes someone’s roadmap. - **Design and ship agent-powered features** such as root-cause analysis, incident triage, and monitor generation—then integrate them into the platform. - **Build eval infrastructure** to make features safe to change: golden datasets, regression suites, offline/online scoring, and the judgment calls for what “good” means. - **Own retrieval and context pipelines** over customer metadata, lineage, and query history. - **Instrument agent behavior in production** (traces, failure taxonomies, cost/latency budgets) to close the loop on quality. - **Partner with data science** on detection quality and experiment design; partner with PM on what an agent should do vs. what it merely can do. - **Set the technical bar for building with LLMs**: patterns, guardrails, and internal tooling other engineers reuse. ### What We’re Looking For - **Built agents in production** with real autonomy and internal loops (not just framework integration). Agents should use tools and decide what to do next without a human in the middle—and you kept them running once real users arrived. - RAG with a wrapper doesn’t count. - A set of MCP tools pointed at an API doesn’t count. - **Run evals and monitor agents after launch**. Since agents are non-deterministic, you’ve owned an eval framework (golden datasets, regression suites, offline/online scoring) and watched agents in production. - **Python + ML/data science background**. Python is your daily language; you’re solid on backend work (distributed systems expertise not required, but you understand model behavior). - **Work from a problem, not a spec**: design experiments, build the smallest version to test, and take what works into production. - **Use AI tools daily** (e.g., Claude or equivalents) as part of your workflow. - **Prefer shipping over polishing**: you can identify which problems need depth and say no to approaches that demo well but fail in production. **Nice to have** - Statistics and hypothesis testing (applied) - Building and maintaining MCP servers - Experience in data/cloud (Snowflake, Databricks, dbt, Airflow) ### This Is Not For You If - Your AI work is retrieval with a wrapper, or MCP tools pointed at an API—nothing that decides and acts on its own. - Your LLM experience is prototypes/notebooks/demos that never carried production traffic. - You need a fully specified problem before you start, or you’re uncomfortable with ambiguity. ### Why Monte Carlo - You’ll build where the market is forming, not where it’s settled. - Series D, **$236M raised** (backed by Accel, Redpoint, Notable Capital, ICONIQ Growth, Salesforce Ventures). - Customers include HubS

Listing freshness

CronJobs last confirmed this listing 18h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord