Apply Edge Start your job search

ML Researcher, Speech

DeepRec.ai · San Francisco, CA

Apply & track with Apply Edge

Machine Learning Researcher, Audio $250,000 – 300,000+, Equity + Bonus Remote (US & Europe) / San Francisco, CA (Hybrid preferred) Full-time / PermanentDeepRec has partnered with a fast-growing, revenue-generating voice AI company empowering enterprises to build AI phone agents at scale. Recent Series C funding with backing from leading Silicon Valley investors, they are building the models and infrastructure that make voice the primary interface between businesses and their customers.This company has built all of their current models completely in-house, and every model ships to real, paying customers almost immediately. No speculative research track here. If you want your work to hit production within weeks, not years, this is that role. The OpportunityThe research team are working toward a single, ambitious goal: a fully speech-to-speech conversational AI model that understands and responds like a human, in real time. You'll work across the core building blocks of that roadmap, such as: speech-to-text, text-to-speech, neural audio codecs, and getting LLMs to understand and reason over audio directly.You'll take ideas from theory through large-scale training to production inference serving millions of calls a day, working closely with engineering and product teams to get your research into real customer environments fast. What You'll DoBuild TTS models that sound natural, expressive, and humanBuild STT systems that stay accurate with accents, background noise, and messy phone linesWork on neural audio codecs, compressing audio efficiently without losing qualityExplore LLM-Audio UnderstandingPrepare and manage large audio datasets, and design how models are trained on themRun training across many GPUs at once, keeping an eye on cost and speedRun fast, well-designed experiments to test what actually worksMake sure models run fast and reliably in production, not just in benchmarksWhat You'll Bring EssentialA genuine, self-driven interest in this area of research, shown through your own projects or papers, not just what a past employer asked you to doHands-on experience building or improving TTS, ASR, speech-to-speech, or neural audio codec systemsA track record of original, hands-on technical work, rather than off-the-shelf tools or common tutorial-style projectsStrong Python skills and experience training or running large modelsComfortable working autonomously DesirableExperience getting LLMs to understand or reason about audio - the team's top priority right nowExperience with model distillationExperience fine-tuning or training language or speech models with reinforcement learningPublished research or open-source work in speech or language AIBackground working with real-time speech systems or phone-based systemsA PhD is welcome but not required. Strong, independent work matters more than qualificationsWe encourage you to apply even if you don't meet every requirement. The right mindset and genuine curiosity matter as much as the resume. What's In It For YouGround-floor seat in a lean and growing research team with real scope to shape it as it scalesEvery research output ships to real customers, no long speculative research projectsHigh autonomy to shape your own research direction, tooling, and (for senior hires) the team itselfWork across ASR, TTS, neural codecs, and the frontier of LLM-audio understanding & speech-to-speech modellingWell-capitalised, fast-moving environment without big-lab bureaucracyHealthcare, dental, vision, meaningful equity, and every tool you need to succeedRemote-friendly across the US