Senior Research Engineer
Nxt Level · New York City Metropolitan Area
Apply & track with Apply EdgeResearch Engineer, Clinical Reasoning
This gives the team access to a unique feedback loop: real patient interactions, real clinical workflows, and real-world healthcare data that can be used to test, evaluate, and improve clinical AI systems over time.The mission is to make high-quality healthcare more accessible through AI while continuously improving clinical safety, accuracy, reasoning quality, and trust.About the RoleOur client is hiring a Research Engineer, Clinical Reasoning to help build the reasoning systems, learning methods, retrieval algorithms, post-training data pipelines, and evaluation infrastructure behind its clinical AI platform.This role blends research and engineering. The right person can run experiments, post-train models, build high-quality data pipelines, debug ML systems, and ship production-quality infrastructure that helps clinical AI improve with every iteration.This is not a pure research role and it is not a generic software engineering role. It is designed for someone with a strong spike in ML or AI research and the engineering ability to turn open-ended research ideas into working systems.What You’ll BuildAgentic Clinical ReasoningDesign and implement next-generation clinical AI reasoning systemsBuild architectures for reasoning, reflection, verification, tool use, routing, uncertainty handling, and escalationCreate systems where specialized agents and models work together to support safe, reliable clinical decisionsHelp clinical AI systems earn greater autonomy over time through better reasoning, evaluation, and feedback loopsEvaluation and MeasurementBuild evaluation platforms, rubrics, simulations, and experiments that measure clinical AI performanceIdentify when a benchmark score rewards the wrong behaviorDetermine whether improvements should come from reasoning, retrieval, model behavior, data, or engineering changesMeasure whether each intervention genuinely improves safety, accuracy, usefulness, and clinical reliabilityModels, Learning, and Post-TrainingRun post-training experiments to improve clinical AI behaviorApply methods such as fine-tuning, distillation, reinforcement learning, preference optimization, and prompt or system optimizationBuild training data, feedback, reward, and experimentation pipelinesPartner with in-house researchers to translate training objectives into concrete data and evaluation specificationsScale training infrastructure and experimentation workflowsCreate and manage both real-world and synthetic data pipelinesSearch, Retrieval, and GroundingBuild search, ranking, retrieval, and grounding algorithms that connect clinical reasoning to trusted medical evidence, patient context, and partner-specific contentImprove retrieval based on downstream clinical decision quality, not just document relevanceBuild systems that help clinical AI produce grounded, evidence-backed, trustworthy answersWhat We’re Looking ForStrong spike in ML or AI research, especially in areas like post-training, reinforcement learning, evaluations, interpretability, model behavior, or agentic systemsStrong software engineering ability with the ability to build reliable tools, debug systems, and ship working infrastructureExperience post-training models and building or curating high-quality post-training dataAbility to design, evaluate, and improve data pipelines for model training and clinical reasoningStrong judgment around data quality, task realism, evaluation design, and model behaviorFast problem-solving and code comprehension skills, especially in unfamiliar codebases or ML pipelinesComfort debugging PyTorch workflows, model training issues, data pipelines, and evaluation systemsAbility to translate research goals into concrete datasets, experiments, and measurement frameworksAbility to work across research and engineering, turning open-ended AI problems into usable systemsClear communication with research, engineering, clinical, product, and data teamsIdeal BackgroundSuccessful candidates may come from backgrounds such as:Applied ML researchResearch engineeringPost-training or reinforcement learningModel evaluation and interpretabilityData science with strong engineering depthSoftware engineering with a research-oriented AI focusClinical AI, healthcare AI, or safety-critical AI systemsThe team cares more about depth of work, research judgment, engineering ability, and ability to ship than a specific credential or career path.Bonus ExperiencePublished research, patents, meaningful open-source contributions, or novel production ML systemsExperience building AI systems at an early-stage or high-growth companyExperience in healthcare, clinical AI, or another regulated or safety-critical domainFamiliarity with clinical workflows, healthcare data, HIPAA, FHIR, EHR systems, or HL7Experience with human-feedback systems, RLHF, simulation, or synthetic data generationExperience in AI safety, bias detection, calibration, fairness, or model reliabilityExperience building evaluation harnesses for LLMs, agents, or clinical decision systemsWhy This OpportunityJoin a frontier AI lab focused on clinical reasoningWork with real patient interactions and real clinical workflowsBuild systems that help clinical AI reason, learn, retrieve evidence, and improve over timeOwn problems end-to-end across research, data, training, evaluation, and deploymentPartner with researchers, engineers, clinicians, and product teams on high-impact healthcare AI systemsBuild in a safety-critical domain where evaluation, grounding, and reliability truly matterHelp define how clinical AI earns trust and autonomy over timeIdeal Candidate ProfileThe ideal candidate is a research-minded engineer with a real spike in ML, AI research, post-training, reinforcement learning, evaluations, or model behavior.They can understand research objectives, build the data and evaluation systems behind them, debug ML pipelines, post-train models, and ship production-quality infrastructure. They have strong data taste, strong engineering ability, and the judgment to build clinical AI systems that are measurable, grounded, and safe.This person wants to work at the intersection of frontier AI research and real-world healthcare.