Machine Learning Researcher, Audio
Bland · San Francisco
About this role
## Machine Learning Researcher, Audio **Location:** San Francisco, CA or Remote ### About Bland At Bland.com, our mission is to empower enterprises to build AI phone agents at scale. We’re a fast-growing team reimagining how customers interact with businesses through voice—building the models and infrastructure that make those interactions feel natural, reliable, and genuinely human. Bland has raised **$100M** from leading Silicon Valley investors. ### The Role As a **Machine Learning Researcher (Audio)**, you’ll work on foundational research and development across Bland’s voice stack: - **Speech-to-Text (ASR)** - **Large Language Models** - **Neural Audio Codecs** - **Text-to-Speech (TTS)** This is not a narrow research role. You’ll take ideas from theory to large-scale training and production inference systems serving millions of calls per day. You’ll design new modeling approaches, validate them with rigorous experimentation, and collaborate with engineering teams to deploy in real customer environments. ### What You Will Do **Build and Scale Next-Generation TTS Systems** - Design and train large-scale text-to-speech models for expressive, controllable, human-sounding output - Develop neural audio codec-based TTS architectures for efficient, high-fidelity generation - Improve prosody modeling, question inflection, emotional expression, and multi-speaker robustness - Optimize for real-time, low-latency inference in production **Advance Speech-to-Text Modeling** - Build and fine-tune large-scale ASR systems robust to accents, noise, telephony artifacts, and code switching - Leverage self-supervised pretraining and large-scale weak supervision - Improve transcription accuracy for real-world enterprise scenarios, including structured extraction and conversational nuance **Pioneer Neural Audio Codecs** - Research and implement neural audio codecs with extreme compression and minimal perceptual loss - Explore discrete and continuous latent representations for scalable speech modeling - Design codec architectures that enable downstream generative modeling and controllable synthesis **Develop Scalable Training Pipelines** - Curate and process massive audio datasets across languages, speakers, and environments - Design staged training curricula and data filtering strategies - Scale training across distributed GPU clusters with focus on cost, throughput, and reliability **Run Rigorous Experiments** - Design ablation studies to isolate architectural impact - Measure improvements using objective metrics and perceptual evaluations - Validate ideas quickly through focused experiments ### What Makes You a Great Fit - **Deep Research Foundations:** self-supervised learning, multimodal modeling, or generative modeling; ability to derive new formulations and implement efficiently - **Expertise in Voice Modeling:** hands-on experience scaling TTS/STT/codec systems; strong intuition for audio quality, prosody, and conversational dynamics - **Systems and Hardware Awareness:** experience training/serving large models on modern accelerators; inference optimization (quantization, kernel optimization, memory efficiency); understanding of real-time constraints - **Experimental Rigor:** controlled experiments, meaningful ablations, comfort with offline benchmarks and live production metrics - **Builder Mentality:** ownership from research through deployment; excited by ambiguous, unsolved problems ### How You Show Up - Treat unsolved problems as opp
Listing freshness
CronJobs last confirmed this listing 1h ago. If its source stops confirming the opening for seven days, this page is removed from active inventory.