CronJobs

ai-ml jobs

Machine Learning Engineer (Evals and Voice Models)

Aircallioinc · San Francisco Office

onsiteunknown$181,000–$181,000Posted Sep 15, 2026machine learningLLMsTTSASRspeech-to-speechLLM-as-judgesynthetic dataRAG evaluation

Apply on the employer site

About this role

## About Aircall Aircall is an AI-powered customer communications platform used by 22,000+ companies worldwide to drive revenue, resolve issues faster, and scale customer-facing teams. We bring voice, SMS, WhatsApp, and AI into one seamless workspace. ## Role: Machine Learning Engineer (Evals and Voice Models) We’re looking for someone to build out the evaluation foundation across our AI Voice and Messaging products (and other agentic products). You’ll work on voice models, agent capability evals, benchmark design, LLM-as-judge systems, failure analysis, and the infrastructure that ties it all together—establishing shared metrics, test sets, and tooling to measure accuracy, resolution quality, and safety consistently across products. You’ll also set up repeatable pipelines for regression testing and benchmarking as models and features evolve, so teams can ship confidently without reinventing evaluation methodology for each product. ## Key Responsibilities - Design and document comprehensive evaluation frameworks for Aircall’s AI agents across voice, chat, and messaging. - Train and fine-tune voice models (TTS, ASR, speech-to-speech) using production and synthetic data; iterate on architecture, data mix, and training strategy to improve accuracy, naturalness, and latency. - Assess AI-generated solutions across training pipelines, experimentation setups, debugging processes, and optimization strategies. - Analyze system design decisions to identify strengths, weaknesses, and potential failure points. - Create annotation guidelines and workflows for human-labeled evaluation data; calibrate LLM-as-judge systems against human raters to keep automated evals trustworthy over time. - Build and maintain live quality monitoring for deployed AI agents, tracking accuracy, resolution rate, and safety signals in production; flag model/data drift before it impacts customers. - Own the metric contract for every published AI metric (definition, population, grain, rollup, validity window). - Build release gates: an offline regression suite each AI surface must pass before a prompt, model, or config change ships—measuring reliability across repeated trials, not just average pass rates. - Build voice-specific evaluation using simulated callers across accents, languages, background noise, barge-in, DTMF, and tool failures; treat latency and ASR accuracy as first-class quality metrics. ## Minimum Qualifications - BS in Computer Science, Machine Learning, Statistics, or related field - 3+ years of experience in ML Engineering or Applied ML (8+ years overall) - Strong experience evaluating supervised, unsupervised, LLMs, and deep learning models - Hands-on experience in failure analysis and evaluating LLMs - Experience building automated evaluation systems - Strong communication skills to explain complex technical concepts to technical and non-technical audiences - Hands-on experience training or fine-tuning voice/speech models (TTS, ASR, or speech-to-speech), including data pipeline construction and experimentation ## Preferred Qualifications - MS / PhD in Computer Science, Machine Learning, Statistics, or related field - Experience evaluating LLMs or agentic systems (e.g., LLM-as-a-judge, RAG evaluation) - Experience with synthetic data generation and prompt engineering - Experience training/fine-tuning voice models at scale, including synthetic data generation, model distillation, or low-latency inference optimization for production voice agents ## Co

Listing freshness

CronJobs last confirmed this listing 3h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.

Browse all software engineering jobs →

Follow fresh jobs in Discord