Senior Researcher, Evals
techire ai · San Francisco Bay Area
Apply & track with Apply EdgeAbout the jobHow do you actually evaluate an AI model when accuracy on a benchmark only tells part of the story?A well-funded AI company developing next-generation foundation models is looking for a Senior Research, Evals to rethink how model performance is measured across real-world interactions.The roleYou’ll work on evaluation frameworks that go beyond static benchmarks, looking at how models reason, remember, adapt and interact over time.This is a hands-on research role sitting between model development and evaluation. You’ll help define what should be measured, work out how to measure it, then build the systems that bring those evaluations directly into model development.What you'll doDevelop new evaluations for reasoning, memory, interaction and model behaviourBuild evaluation pipelines with robust statistical analysisCreate quantitative metrics for qualities that can be difficult to measure objectivelyIntegrate evals directly into model training and research workflowsDesign user studies and behavioural experiments to understand real-world model performanceWhat you'll bringExperience building evaluation frameworks for generative models across text, audio or multimodal AIStrong technical and analytical skillsExperience turning open-ended research questions into working evaluation systemsA good understanding of metrics, experimentation and statistical analysisAn interest in measuring subjective qualities such as naturalness, adaptability and interaction qualityExperience with alignment or model behaviour research would be useful, but isn’t essential.You’ll work closely with researchers building new foundation models, helping understand whether changes are creating meaningful improvements rather than simply moving benchmark scores.Base salary is $200k-$350k DOE + generous equity.Based in San Francisco, New York or London, working hybrid.If you’re interested in researching better ways to measure how advanced AI systems actually behave, we’d love to hear from you.All applicants will receive a response.