Applied Research Scientist - GenAI
Galtea · Greater Barcelona Metropolitan Area
Apply & track with Apply EdgeAbout GalteaGaltea is building the evaluation platform for AI teams. We help companies ship better AI, faster, by making it easy to test, evaluate, and monitor AI agents and LLM-powered products at scale. We're backed by top-tier investors and working with some of the most sophisticated AI teams in Europe, including major financial institutions and tech companies.We're a small, high-conviction team. Every hire matters.The roleAn AI system can only improve as fast as it can tell whether it got better. Capability is cheap now. Judgement is not. The limit on every self-improving system is the quality of the signal that grades it, and that signal is still the weakest part of the stack.Galtea is seeking an Applied Research Scientist to work on that limit. You will bring frontier research into what we ship, publish the results that stand on their own, and design the direction of our core technology alongside the founders. This is the first fully dedicated research seat in the company, and it sits at the centre of the product rather than beside it.What you'll do - Three things, in roughly equal measure.— Bring the frontier into the product. Read the field, judge what is real, and turn what is real into capability our customers can use. You design the experiments that say whether a new method beats what we run today, train and distil models where training is the right answer, and take the result to production. Research that stays in a notebook doesn't count. Neither does shipping something we can't measure.— Do original research and publish it. Some of the hardest questions in this field have no answer yet, and we hit them in customer work every week. You pick the ones worth solving, solve them properly, and write them up. Papers, technical reports and open-source releases are part of the job, not a side project you have to negotiate for.— Shape where the technology goes. You work with the founders on what Galtea should be able to do in two years that nobody can do today, and you turn that answer into a research agenda and then into product. Over time you own the R&D half of the product backlog and build the team that grows around it.What we're looking forA researcher who ships. You can design an experiment and defend its statistics, and you can put the resulting model behind an endpoint customers depend on. Most people do one of those.Must have— Research depth. PhD in Computer Science, Machine Learning, NLP or a related field, or an equivalent record of original research in industry. You read frontier literature, you can tell a real result from a well-marketed one, and you can reproduce what matters.— Production track record. You have taken models or AI services to production and kept them alive there. You know what changes between a result and a service under load, and you have owned that transition yourself instead of handing it over.— Post-training expertise. Hands-on command of fine-tuning, distillation and preference optimization for language models, with the open-source stack around them: PyTorch, Hugging Face, modern inference servers, parameter-efficient methods, quantization, and GPU training at the scale a focused team can afford.— Evaluation instinct. You understand how automated evaluation fails. Model-graded evaluation, human annotation, agreement statistics, calibration, and the ways a grader can be confidently and systematically wrong are familiar territory, or you are visibly hungry to make them so.— Experimental rigour. Fair baselines, ablations that isolate the cause, sound statistics, and the honesty to report a negative result as a result.— Engineering strength. Strong Python, and the judgement to write research code another engineer can take to production. You work in a modern AI-assisted development workflow without letting review standards or architecture slip.— Autonomy. You are our first dedicated research hire. You can turn a company goal into a research agenda, sequence it, and report progress honestly with no research organisation around you.— Communication. You can explain a method to an engineer, a result to a customer and a trade-off to a founder, in writing and out loud.Nice to have— Experience training or distilling grader, reward or classifier models.— Experience with reinforcement learning or preference optimization.— Experience deploying models into on-premise, air-gapped or otherwise restricted environments, including inference optimization on hardware you don't control.— First-author publications at NeurIPS, ICML, ICLR, ACL or EMNLP, or significant contributions to widely used open-source AI projects.— Experience evaluating agentic or multi-turn systems, including tool use and retrieval.— Experience with synthetic data generation and its quality control.— Multilingual work, in particular Spanish or Catalan.Education and languages— PhD in a relevant technical field, or a Master's degree with an equivalent research and production record.— Fluent written and spoken English. Spanish is a plus.What we offer— The first dedicated research seat in the company, with the founders as your counterparts on direction.— Direct exposure to the most sophisticated AI engineering teams in Europe.— Competitive salary + equity (ESOP).— Hybrid setup in Barcelona.— Compute for your experiments, and real support to publish and present what comes out of them.— A learning and development budget for courses, certifications and conferences.— Paid AI tooling and model credits, and the expectation that you use them.— A small, high-ownership team where your work has real impact from day one.