Agent Evaluation & Instrumentation Engineer
Umanist NA · Pune City, Maharashtra, India
Apply & track with Apply Edge🚀 Agent Evaluation & Instrumentation Engineer
Location: Pune | Experience: 8–11 Years | CTC: Up to ₹33 LPA(576)Notice Period: Immediate–45 Days | Preference: Local Pune CandidatesInterview: 2 Technical Rounds + Client Round + HR🔥 MANDATE CRITERIA – NON-NEGOTIABLE6+ years in ML/Data/Software Engineering with strong evaluation/quality focus Hands-on LLM/ML evaluation experience Strong Python skills Experience with Promptfoo / DeepEval / custom evaluation frameworks Strong understanding of LLM/Agent evaluation design & statistical rigour Hands-on data analysis & metric interpretation Experience with tracing & observability tools Strong technical documentation, analytical & stakeholder-management skills Understanding of Telco customer intents & customer journeys 2+ years stability in an organization 💼 Key ResponsibilitiesDesign and maintain LLM/Agent evaluation suites, golden sets, regression packs & adversarial tests Build continuous evaluation & scoring pipelines Conduct quality/operability gate reviews and reproduce evaluation results Analyse model regression, failure modes, defects & evaluation trends Calibrate quality thresholds, judges & evaluation datasets Monitor evaluation drift and maintain hold-out/golden sets Prepare technical quality reports & evaluation findings Collaborate with AI/ML, Engineering & Operations teams on quality improvements Mentor team members and communicate findings to technical/client stakeholders ⭐ GOOD TO HAVEAgentic AI | RAG | Telco AI | AI Red-Teaming | Responsible AI & Safety | Judge Calibration | Management PresentationsIdeal Profile: AI/ML Evaluation + Python + Data Analysis + LLM/Agent Quality with strong communication and client-facing skills.Skills: deepeval,quality thresholds,llm/agent evaluation design,python,llm/ml evaluation,telco customer intents,regression packs,technical documentation,statistical rigour,tracing,model regression,evaluation drift,promptfoo,ml/data/software engineering,golden sets,failure modes