Senior | Staff Software Engineer - AI / ML
Snorkel AI · San Francisco, CA (Hybrid)
About this role
## About Snorkel At Snorkel, we believe meaningful AI doesn’t start with the model—it starts with the data. We help enterprises transform expert knowledge into specialized AI at scale, and we’re scaling our engineering and research teams following our **$350M Series E** (September 2026). ## The Role Frontier AI data is expensive to make and hard to measure. You’ll help make the evaluation process **faster, cheaper, and more rigorous** using ML and AI. You’ll be an early member of **ML & Research Engineering**, studying how frontier-grade data is generated and evaluated, forming and validating hypotheses against real production data, and shipping the winners at scale—helping shape the discipline’s direction, standards, and team. ## What you’ll work on - **Efficient agentic evals:** Reduce the cost of long-horizon agent evaluation with adaptive sampling, statistically grounded early stopping, model cascades, caching, and cheap-first gating. - **AI model routing:** Route every eval/judge call to the cheapest model that clears the quality bar, with fallback, monitoring, and cost attribution. - **Fine-tuned small models:** Fine-tune and serve open-weight models (e.g., LoRA / parameter-efficient methods) when they match frontier quality—and know when they don’t. - **Predictive difficulty:** Build models that estimate task difficulty for frontier systems before running rollouts. - **Measurement for AI data:** Create golden datasets; quantify accuracy and calibration of LLM-as-judge systems; make quality reproducible across projects. - **Research to production:** Turn prototypes into reusable, configurable components used by deployed engineers and researchers. ## What you’ll bring - **5+ years** building production ML or software systems, with end-to-end ownership from prototype to production - Hands-on experience running **LLM/ML workloads in production**; comfort reasoning about non-deterministic systems - Strong grounding in **statistics and experimentation** (experiment design, hypothesis testing, sampling, confidence intervals) - Strong **Python and software engineering fundamentals** (testing, code review, system design) - Experience designing evaluations and interpreting results rigorously - A habit of finding high-impact problems early, plus clear communication across researchers, engineers, and business partners ## Nice to have - Fine-tuning/serving open-weight models; knowing when smaller models meet the quality bar - Building LLM evaluation/experimentation platforms, model gateways, or routing systems - Experience with agentic workloads, benchmarks, or RL environments - Track record taking research into production (publications, open-source, or shipped features) - **MS or PhD** in CS, ML, Statistics, or related field ## Why join now - **Frontier problems:** Measure and shape tasks designed to challenge the strongest models - **All AI, no plumbing:** Open ML/LLM problems across the team - **Founding impact:** Help define what ML engineering means at Snorkel and grow the team - **Visible results:** Direct impact on speed, quality, and cost of the AI data frontier ## Compensation (Pay Transparency) Actual compensation will be determined based on skills, qualifications, experience, and geographic location. **Salary range: $208,000–$315,000 USD.** ## Equal Opportunity Snorkel AI is an Equal Employment Opportunity employer and is committed to building a diverse team. Reasonable accommodations are available for applicants with disab
Listing freshness
CronJobs last confirmed this listing 3h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.