Research Scientist
SoTalent · Boston, MA
Apply & track with Apply EdgeResearch Scientist – AI/ML (Foundation Models & Generative AI)IndustryArtificial Intelligence / Machine Learning / Big Tech / Applied ResearchWork SettingResearch & Development environment | Hybrid or on-site | High-collaboration engineering + science teamRole OverviewA research-focused AI/ML role centred on the development, training, and optimisation of large-scale foundation models. The position involves advancing generative AI systems, improving model performance, and translating cutting-edge research into scalable production-ready solutions.Key ResponsibilitiesFoundation Model ResearchDesign and develop large-scale machine learning and foundation modelsResearch improvements in architecture, training efficiency, and model performanceWork on generative AI systems including LLMs and multimodal modelsModel Development & ExperimentationBuild and run large-scale experiments for model training and evaluationDevelop novel algorithms for representation learning and optimisationAnalyse model behaviour, performance, and failure modesData & Training PipelinesDesign datasets and data strategies for model pretraining and fine-tuningWork with large-scale distributed training systemsImprove data quality, filtering, and augmentation methodsEngineering & ImplementationCollaborate with ML engineers to scale research prototypes into production systemsOptimise models for inference efficiency, latency, and costUse frameworks such as PyTorch, TensorFlow, or JAXCollaboration & Research OutputWork closely with applied scientists, engineers, and product teamsPublish research findings in top-tier ML conferences (optional depending on org)Contribute to internal research direction and technical strategyRequirementsEducationPhD (preferred) or Master’s in Computer Science, Machine Learning, AI, Mathematics, or related fieldExperience2–5+ years experience in ML research or applied AI (varies by level)Strong background in deep learning and neural networksExperience with large-scale model training or distributed systemsTrack record of building or researching transformer-based architectures or similarTechnical SkillsStrong proficiency in PythonExperience with PyTorch, TensorFlow, or JAXKnowledge of transformers, LLMs, or diffusion modelsUnderstanding of optimization, GPU training, and scaling ML systemsExperience with distributed computing or high-performance ML infrastructureCore CompetenciesStrong research and experimental design mindsetAbility to translate theory into working systemsAnalytical thinking and model debugging skillsCollaboration across research and engineering teamsComfort working in fast-moving, ambiguous R&D environmentsRole FocusFoundation model developmentGenerative AI innovationLarge-scale ML experimentationResearch-to-production AI systemsAdvanced deep learning architecture design