Senior Data Scientist AI Evaluation
Alpaca · Remote - Americas
About this role
**Senior Data Scientist, AI Evaluation** **Who We Are** Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more. We’re a licensed financial services company serving hundreds of financial institutions across 40 countries with institutional-grade APIs—supporting broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, totaling over 10 million brokerage accounts. **Your Role** We’re looking for a Senior Data Scientist, AI Evaluation to design how Alpaca measures whether our models and agents are actually right. You’ll turn ambiguous quality questions into ground truth, scoring methods, and eval loops the company can trust—then use those results to improve the systems. You’ll build on an established data foundation, focusing on raising quality and speeding up safe rollout, and you’ll own the quality bar independent of the teams that build and optimize these systems. You care just as much about whether an answer is correct as you do about whether a model can generate one. You’ll partner with Product, Engineering, Analytics Engineering, and business stakeholders to define what “good” looks like, build the evaluations that test it, and close the loop so evals drive iteration. **What You’ll Do** - **Design AI evaluations:** Define ground truth, metrics, and scoring methods for models and agents. - **Build repeatable eval loops:** Track quality over time and catch regressions before release. - **Partner on infrastructure:** Work with engineering and analytics engineering to operationalize eval harnesses. - **Drive iteration:** Translate eval results into actionable recommendations for system improvements. - **Establish quality standards:** Set evaluation guidelines, documentation, and review practices. - **Mentor and align:** Foster evaluation best practices and build a culture of measurable AI quality across the team. **What We’re Looking For** - Track record of **quantitative measurement rigor** (e.g., LLM/model evaluation, metric validation, or experimentation). - Strong **statistical and ML foundation**—treat evaluations as experiments (sample sizing, confidence intervals, significance, handling non-determinism) and validate automated graders against human ground truth. - Proficiency in **Python and SQL**, with experience evaluating models in production environments. - Strong judgment in defining **quality metrics** and **ground truth** for ambiguous outputs. - Excellent communication and cross-functional collaboration skills. - Strong problem-solving ability in fast-paced, greenfield environments. - **6–10 years** in quantitative data science or ML, with focused experience in measurement or evaluation (a quantitative degree is a plus; equivalent experience is welcome). **Nice to Have** - Hands-on **LLM/agent evaluation in production**, including eval harnesses, LLM-as-judge calibration, and CI regression gates. - Experience evaluating **text-to-SQL**, analytics agents, or other systems where correctness is verifiable against data. - Background in **fintech/brokerage** (where wrong answers have real business or risk consequences). - Fluency with AI tools in research and engineering workflows. **How We Take Care of You** - Competitive salary & stock options - Health benefits - New hire home-office setup: **one-time USD $500** - Monthly stipend: **USD $150/month** via a Brex Card - Equal opportunity workpla
Listing freshness
CronJobs last confirmed this listing 2h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.