Founding Research Engineer (SaaS)
ixigo · New Delhi, Delhi, India
Apply & track with Apply EdgeResearch Engineer — Agent Intelligence & EvaluationVoice agents fail in ways traditional software doesn't. A tool call misfires when ASR confidence drops on a regional accent. An LLM hallucinates a policy because upstream latency broke turn-taking. Customer support teams roll these agents back within a week of shipping, and nobody can explain what actually went wrong.We're building self-healing voice agents for enterprise customer support. The system has to know when it's failing, why it's failing, and how to fix itself before a human notices. This role is about building the intelligence layer behind that work: evaluations that catch failure modes before deployment, observability that traces problems across the audio, reasoning, and tool-use pipeline, and feedback loops that let agents learn from their own conversations.What you'll work onThree connected problems.Evaluation frameworks for voice agents. Text-only evals miss most of what matters here (barge-in, prosody, latency-induced errors, cross-turn context loss). You'll design audio-native metrics, generate adversarial conversational datasets across accents and edge cases, and build LLM-as-judge rubrics for task completion, empathy, and recovery from tool failures.End-to-end observability. Tracing a failed customer interaction means correlating audio packets, STT hypotheses, LLM reasoning traces, tool calls, and TTS output back to a single conversation ID. You'll help design the schema and analysis layer that makes cascade failures visible and diagnosable across the stack.Self-improvement systems. Once you can measure and trace, the interesting work is closing the loop: mining production traces for failure patterns, automatically generating targeted fine-tuning data or prompt updates, and validating that fixes hold under adversarial replay.Who we're looking forSomeone who cares about the research questions for their own sake. Papers in Interspeech, ACL, NeurIPS, or EMNLP on speech, dialogue systems, agent evaluation, or human-AI interaction are directly relevant, and we want someone actively tracking the literature.Someone who also cares whether the research ships. Enterprise customer support has real users on the other side of every call, and elegant methods that fall over in production traffic don't help us. If you've felt the gap between a benchmark number and a real-world win, and it bothered you, you'll fit here.Comfortable in Python. Familiar with at least one of the following: speech models (Whisper, Conformer variants and their descendants), LLM tool-use and agent frameworks, or trace and observability stacks (OpenTelemetry, Langfuse, Arize, Hamming, etc). You don't need all three, but you should be ready to pick up the other two once you're in.A current PhD student in ML, NLP, speech, or a related area is the strong default. Exceptional MS students or research engineers with a publication track record are welcome to apply.Nice to havePrior work on evaluation methodology, dataset synthesis, or interpretability. Experience with real-time systems, telephony infrastructure, or streaming pipelines. A blog, a repo, a workshop paper, or anything else that shows how you think about problems in public.What you'll getA small engineering team where the work reaches production, with publication as an expected output of the role. Access to a stream of real enterprise conversation data under proper governance. Mentorship on both the research side and the shipping side. Co-authorship on papers that come out of the work, and the ability to point at a system in production that runs on top of them.