أبلاي إيدج ابدأ البحث عن عمل

AI/ML Data Scientist

Medly AI · London Area, United Kingdom

قدّم وتابع مع أبلاي إيدج
About Medly AIWe're the fastest growing EdTech startup in London, on a mission to change education forever.This is a moment where AI-native companies are reshaping entire industries, from Cursor in code, to Midjourney in image generation, to Harvey in law. Medly is building that company for education.Since launch, we've reached over 400,000 students with retention rates that beat Duolingo, and we've delivered personalised learning with real results. We're backed by impact-focused VCs and supported by partners including UCL, Innovate UK, Microsoft, and Google.The roleWe're looking for an AI/ML Data Scientist to join our London team. You'll work directly alongside our founders and engineers who've previously built at Meta, Atlassian, and more, and you'll own how we measure and improve Medly's AI models.This is the person who tells us the truth about our models. Hundreds of thousands of students rely on our AI to tutor them, mark their work, and generate content pitched at the right level for their exam board. Right now, judging whether a change made that better or worse is the hardest problem we have. You'll build and improve the benchmarks, evals, and experiments that answer it, and you'll be the reason we can ship quickly without breaking what works.You'll have real ownership from day one. This is not a role where you'll be handed a spec.What you'll do- Own our AI benchmarking: help improve and design subject-specific evals that measure tutoring, marking, and content generation quality against what a good teacher would actually say- Improved the eval harnesses and offline test suites that let us ship model and prompt changes with confidence, and catch regressions before students do- Run experiments and A/B tests on live AI features, and make the call on what the results actually mean- Build and curate high-quality datasets for fine-tuning and evaluation, including designing labelling schemes and rubrics that other people can apply consistently- Investigate model failures end to end: find where the AI falls down on notation, working, partial credit, or a specific exam board, and turn that into a fix- Work with our learning developers and teachers to turn pedagogical judgement into something measurable- Get hands-on with our product data: spot patterns, surface insights, and shape what we build next- Contribute to the wider codebase where it makes sense (Python, Postgres, AWS, some Next.js)What we're looking for- At least 2 years in a data science, ML, or applied research role, post-graduation- Hands-on experience evaluating or benchmarking AI/ML systems: you've built evals, designed metrics, or run structured model comparisons, and you can talk about where your own metrics misled you- Strong Python for data work (pandas, numpy, notebooks, and the ability to write code other people can run)- Confident with SQL and relational databases- Solid grounding in statistics and experiment design: sampling, significance, confidence intervals, and knowing when a result doesn't mean what it looks like- Working knowledge of LLMs and how they're trained, evaluated, prompted, and fine-tuned- A degree in Computer Science, Maths, Physics, Statistics, Engineering, Data Science, or a similarly quantitative field- Genuine curiosity about how AI systems behave in the wild, not just on a leaderboardBonus points- Experience with eval frameworks and LLM-as-judge setups, including their failure modes- Fine-tuning, RLHF, DPO, or preference data collection- Built annotation or human-evaluation pipelines, and managed the people doing the labelling- Worked on AI products in production where quality was subjective and hard to measure- Any teaching, tutoring, or exam-marking experience, or familiarity with the UK curriculum (GCSE, A-Level, AQA/Edexcel/OCR)- Published research, open-source contributions, or writing about evaluation- Cloud infrastructure experience (AWS)- Something you've built in your own time that you'd love to show usRight to workYou must have the right to work in the UK. We're not able to offer visa sponsorship for this role.Who thrives here- You care about quality and take pride in work that holds up under scrutiny- You're comfortable being the person who says the model got worse, with the evidence to back it- Curious about how things work, and willing to dig into other people's code and data to find out- You take initiative. If something looks off, you flag it or fix it rather than waiting to be told- You can explain a technical result to someone who isn't technical, and be trusted on it- Comfortable in a startup environment. Things move very quickly and priorities shift oftenWhat you'll get- Real ownership of how Medly measures AI quality, on a product used by hundreds of thousands of students- Direct work with founders and senior engineers with experience at Meta, Atlassian, and beyond- Hybrid setup: 4 days in our central London office, one day work from home- A team that takes your ideas seriously and ships them- Room to grow into ownership of Medly's wider AI and data function