Machine Learning Intern
Bland · San Francisco
About this role
**THE ROLE: Machine Learning Research Intern (Audio)** As a Research Intern at Bland, you’ll own a focused research project across our voice stack—**speech-to-text (ASR), large language models, neural audio codecs, or text-to-speech (TTS)**. You’ll work alongside our research team on the same problems they’re tackling (not a side track), scoped around a single meaningful question you can answer within your internship. The goal is a result worth shipping, publishing, or both. Interns here regularly see their work reach production systems handling **millions of calls**. --- ## What You Will Do - **Own a research question end-to-end** - Take one well-scoped problem from literature review through implementation, experimentation, and results - Design **ablations** that isolate what caused an improvement - Present findings to the research team and **defend your methodology** - **Work on real systems** - Train and evaluate models on large-scale, real-world **telephony audio** (accents, noise, artifacts) - Use our **distributed GPU infrastructure** (not toy-scale setups) - Where it makes sense, collaborate with engineers to move results toward **production** - **Choose your depth** (based on your background and interests) - Expressive, controllable **text-to-speech** (prosody & emotion modeling) - **Neural audio codecs** and discrete/continuous speech representations - **ASR robustness** for telephony, accents, and code switching - **Real-time / streaming inference** under latency constraints - Full-duplex conversation and **turn-taking** dynamics --- ## What Makes You a Great Fit - **Research foundations** - Currently pursuing an **MS or PhD** in ML, CS, EE, or related field (or equivalent research experience) - Comfortable reading papers and reimplementing without hand-holding - Experience with **self-supervised, generative, or multimodal** modeling - **Audio or speech grounding** - Hands-on work with speech/audio models (TTS, ASR, codecs, or audio representation learning) - Strong intuition for audio quality and what makes synthetic speech sound wrong - Prior publications or open-source contributions in speech/language AI are a strong signal (not required) - **Engineering ability** - Fluent in **PyTorch** and comfortable working in a real codebase - Able to run experiments on **GPU clusters** independently --- ## How You Show Up - You identify the single experiment that validates an idea in **days, not months** - You measure everything and let data drive decisions - You’re honest about negative results (they help narrow the search) - You’re obsessed with making voice agents sound truly human - You use AI tools aggressively to amplify your impact --- ## Benefits - Competitive intern compensation - Mentorship from researchers working on frontier voice AI - Every tool you need to succeed - Beautiful office in **Levi’s Plaza, SF** with rooftop views - A real shot at a **return offer** --- *Additional Information* Bland is an **equal opportunity employer** and is committed to equal employment opportunities. Bland participates in the **E-Verify** employment verification program. All new hires must complete the **Form I-9**, and employment will be verified through E-Verify as part of onboarding.
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.