Principal Research Scientist | Speech & Voice Foundation Models
Intelix.AI · Berlin, Germany
Apply & track with Apply EdgePrincipal Research Scientist | Speech & Voice Foundation ModelsGermany of DACH region | UK and parts of Europe will be consideredText-to-speech, voice cloning, speech synthesis, real-time conversational voice.$270,000–$500,000 base, plus bonus, equity and benefits (US).Permanent, full-time.The roleBuild the models that are the product. This is full-stack research ownership: you frame the question, run the experiments and ship the result. Research isn't finished until it's in production and measurable.ResponsibilitiesTrain foundation models: pre-training, reinforcement learning, reward modelling, post-training, new architectures, scaling.Build and improve voice and speech models across TTS, STT and speech-to-speech.Design the evaluation that proves the work: benchmarks, eval loops, LLM-as-judge, failure analysis. Evaluation is treated as a research product in its own right, not a pre-launch checkbox.Work on frontier problems adjacent to the roadmap: multimodal, agents and tool use, test-time compute.Take models into production alongside the serving engineering team, inside a real-time latency budget and across a large multilingual footprint.EssentialHands-on foundation-model training: pre-training, RL, reward modelling, post-training, scaling. Fine-tuning or building on top of someone else's model is a different discipline and is not what this role is.Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest signal; TTS and ASR both count. Text-only research does not transfer.Evidence we can point at: papers, shipped models, open-source contributions, or systems running in production.DesirableEvaluation depth: benchmarks, eval loops, quality measurement, failure analysis.Publications at ICML, ICLR, NeurIPS, EMNLP, ACL, AAAI, Interspeech or ICASSP.PhD in ML or NLP, or equivalent practical experience you can point to.Frontier exposure: multimodal, agents, tool use, test-time compute.Public work: side projects, open-source, technical write-ups.Seniority levelMid-Senior levelIndustryArtificial Intelligence · Software DevelopmentEmployment typeFull-timeSkillsSpeech Synthesis (TTS) · Automatic Speech Recognition (ASR) · Foundation Models · Deep Learning · PyTorch · Reinforcement Learning · Multimodal Machine Learning · Model Evaluation · Research · Python