Head of Research
talp · San Francisco Bay Area
Apply & track with Apply EdgeEmployment type: Full time - HybridCompensation: $150,000 – $300,000 · equityAbout TalpTalp is simulation infrastructure for enterprise decisions. We build a company's customer base from its own behavioral data so decisions can be tested before they are made, and we already work with large enterprises across three regions.We are the wind tunnel for decisions. What we install stays and sharpens every time it runs, which is why we call it infrastructure rather than a tool. Testing decisions this way will soon be as ordinary as version control.We have raised $2.2 million to date, most recently at a $25 million valuation, from Formus Capital, Sunshine Lake Ventures, Aito Capital, the a16z Scout Fund, and a group of GPs and founders.The RoleEveryone in this field renders the simulation. Nobody renders the proof. Confidence scores, calibration methods, benchmarks measured against human evidence... None of it has been published by anyone in this category, and the company that publishes it first will define how everybody else gets measured.That is the center of this job.You will lead our scientific direction and turn research into working systems. Your scope spans modeling, evaluation, and the infrastructure that supports experimentation and deployment.This is a hands-on leadership role. You will design experiments, write code, inspect data, and work closely with engineering while building a research team. You should be comfortable asking difficult scientific questions and taking full responsibility for how they are answered.ResponsibilitiesSet the research direction: Identify the most valuable questions, prioritize experiments, and connect scientific progress directly to product improvements.Develop and improve models: Work across data preparation, model adaptation, training, and fine-tuning. Reproduce relevant research, challenge its assumptions, and test approaches against strong alternatives.Define how quality is measured: Design benchmarks and experiments that assess behavioral accuracy, consistency, and usefulness. Compare simulations against human evidence while accounting for uncertainty, population variance, and data limitations.Build reliable evaluation systems: Develop repeatable workflows for comparing models, investigating failures, and detecting regressions. Keep datasets, experiments, and results strictly traceable, protecting evaluation data from leaking into development.Combine human and automated judgment: Create clear scoring criteria, review processes, and automated evaluators, then verify that those evaluators consistently agree with qualified human reviewers.Make research practical to run: Partner with engineering on data pipelines, experiment tracking, training workflows, and model serving to maximize iteration speed, reliability, and compute efficiency.Carry improvements into production: Follow promising results through implementation and deployment, verifying that their advantages hold outside the original experiment.Build the research team and culture: Recruit and mentor researchers and engineers, establish clear standards for rigor, and communicate findings through technical writing, publications, and external collaborations.Who You AreApplied Research Record: A strong track record in applied AI or machine learning research, with evidence of taking ideas from experiments into working production systems.Deep Model Experience: Hands-on experience with language models, modern training/fine-tuning techniques, and rigorous model evaluation.Core Engineering Skills: Strong Python skills and deep familiarity with frameworks such as PyTorch or JAX.Sound Statistical Judgment: You reason intuitively about sampling, uncertainty, and bias, and can tell immediately whether an apparent improvement is statistically meaningful.Systems & Tooling Ownership: Experience building or owning research tooling, evaluation pipelines, or ML infrastructure that other engineers rely on.Team Leadership: Experience leading researchers or substantial research programs, and the judgment to build a team whose strengths complement your own.High Agency: The ability to drive technical work autonomously and make clear decisions even when the empirical evidence is incomplete.Intellectual Honesty: Clear communication and the willingness to discard an approach the moment the data challenges it.Typically five or more years of relevant AI or ML research experience, including at least one year leading research teams or substantial research programs. Relevant academic research counts toward this; we also welcome candidates with exceptional demonstrated impact on a shorter timeline.Valuable Additional ExperienceBehavioral science, computational social science, causal inference, or experimental designSynthetic data generation, human feedback loops, annotation systems, or complex behavioral datasetsMulti-agent systems and simulations involving interacting populationsDistributed training, GPU optimization, or efficient model servingYou do not need equal depth in every area. We look for deep research instincts, strong engineering ability, and the judgment to build a team with complementary strengths. A PhD is welcome but not required.Our ProcessWe prioritize thoughtful conversations and clear examples of past work. Our hiring process is designed to help both sides align on mutual fit, working style, and expectations.Reapplication Policy: To ensure a fair and thorough evaluation for all applicants, Talp observes a 90-day waiting period before reconsidering candidates for the same role.Equal OpportunityTalp is an equal opportunity workplace. We welcome applicants of every background and identity. If you need support or an accommodation at any point in the process, let us know and we will arrange it.