Apply Edge Start your job search

Founding ML Researcher

Morpheus Talent Solutions · New York City Metropolitan Area

Apply & track with Apply Edge
Founding Machine Learning ResearcherWe are building the evaluation and training infrastructure for autonomous healthcare AI. We create high-fidelity training gyms that improve frontier models and agents on real clinical, operational, and administrative workflows, using long-horizon evaluations, expert feedback, verifier-driven tasks, and real-world healthcare data. We are backed by premier VCs, commercially live, connected with frontier labs, and hiring to meet immediate delivery needs.We are looking for an exceptional Founding ML Researcher to help us deliver post-training datasets, evaluations, and gym artifacts (environment + tasks + verifiers + trajectories) to leading AI labs. This role is ideal for someone who wants to work at the frontier of applied ML, RL/post-training, and evaluation, and who can move quickly from ambiguity to concrete artifacts that improve model behavior.What You'll Do:Design and build high-fidelity post-training gyms for clinical and administrative workflows, including tools/API interactions, document flows, and realistic constraints.Translate real-world healthcare workflows into rigorous training and evaluation tasks with clear success criteria (verifiers + rubric-based scoring).Analyze model failures and convert them into targeted datasets, preference data, reward signals, verifier improvements, and task redesign.Work directly on lab-facing deliverables: environment specs, task libraries, baseline evaluations, failure reports, and trajectory/post-training data exports.Collaborate with clinicians, domain experts, and engineers to define what “great” model performance looks like in complex healthcare settings.Help build infrastructure for long-horizon agent evaluation, trajectory generation, regression testing, and reproducible evaluation runs.Move quickly from ambiguous research questions to concrete, shippable artifacts under real timelines.What We’re Looking ForStrong ML research or engineering background, ideally with experience a few of the following areas:Reinforcement learning and/or reward modelingPost-training (SFT, preference data, RLHF/RLAIF)Evaluation and benchmarking for LLMs/agentsAgentic systems (tool use, long-horizon tasks, verifiers)Synthetic data / trajectory generationModel behavior analysis and failure taxonomyAlso looking for:Ability to reason from first principles about model failures, task design, and evaluation quality.Strong coding ability and comfort building research infrastructure quickly.High agency, strong ownership, and ability to operate in a fast-moving startup environment.Strong written communication—you can clearly document tasks, environments, rubrics, model failures, and experimental resultsExperience working with clinical and/or life sciences workflows (flexible)Prior experience building evals, agents, RL environments, or post-training datasets for strong modelsExperience with human expert feedback loops, annotation systems, or QA/QC pipelinesPublications, open-source work, or strong project work in ML, RL, LLMs, agents, or healthcare AI.