CronJobs

ai-ml jobs

Machine Learning Engineer, Core Experimentation

OpenAI · Seattle

hybridunknown$437,000–$485,000Posted Sep 24, 2026PythonGoLLMsretrievalranking/recommendationforecastinganomaly detectioncausal inference

Apply on the employer site

About this role

## About the Team The Statsig team within OpenAI builds experimentation, feature rollout, dynamic configuration, and analytics systems that help OpenAI ship products with speed, safety, and evidence. The platform supports teams across ChatGPT, Codex, model measurement, consumer experiences, business subscriptions, developer products, and shared infrastructure—enabling responsible experimentation, clear measurement, and safe rollouts. ## About the Role OpenAI is seeking a **Machine Learning Engineer** to lead the technical direction for **ML-powered experimentation and insights**. This is an **end-to-end, 0-to-1** role where you will build production systems that learn from privacy-protected product and experimentation data to generate **evidence-backed insights** and help teams decide which ideas are worth testing live. A key challenge is ensuring insights and predictions are **traceable, calibrated, useful, and safe** enough to influence real product decisions. Live experiments remain the source of causal validation. You will design systems that make uncertainty explicit, backtest against historical outcomes, compare predictions with online results, learn from misses, and **abstain when evidence is weak**—while preserving clear review, permission, and approval boundaries. ## What You’ll Do - Set and execute the technical roadmap for **Generative Insights** and **Predictive Experimentation**, from prototypes to production adoption. - Build cross-experiment learning systems to retrieve and synthesize historical experiments, detect recurring effects, segment behavior, reanalyze prior results as methods improve, and generate hypotheses with clear evidence and provenance. - Develop predictive models and simulation workflows to estimate likely impact, affected segments, regression risk, and uncertainty before live experiments. - Create high-quality datasets and feature/retrieval pipelines from exposures, events, metrics, experiment metadata, and replay data—with strong lineage, freshness, privacy, and data-quality controls. - Establish rigorous evaluation via offline benchmarks, backtests, calibration, drift monitoring, prediction-to-outcome comparisons, and explicit failure/abstention behavior. - Turn models into durable product, API, and agent workflows that move from insight → experiment design → approval-gated action → measured learning. - Partner with data science and product teams on experiment design, causal inference, sequential decision-making, variance reduction, and the boundary between prediction and causal evidence. - Build reliable services and intuitive workflows so sophisticated ML capabilities are understandable and useful for high-stakes product decisions. - Provide technical leadership across engineering, product, data science, and research partners—raising the bar for production ML quality across the platform. ## You Might Thrive If You - Have led ambiguous **0-to-1 production ML** products where success was measured by better real-world decisions (not only offline metrics). - Have hands-on experience across the ML lifecycle: dataset design, training/adaptation, evaluation, deployment, monitoring, and iteration. - Have depth in one or more areas such as LLM/retrieval systems, ranking/recommendation, forecasting/anomaly detection, causal ML/experiment analysis, or simulation. - Have strong software engineering fundamentals and can build production systems in **Python** across data, backend, and platform boundaries. - U

Listing freshness

CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord