Data Scientist Intern
M Science · New York City Metropolitan Area
Apply & track with Apply EdgeSummer 2027 Data Scientist Intern – Alternative Data
We are revolutionizing research, discovering new data sets, and pioneering methodologies to provide actionable intelligence. Our research teams have decades of experience working with massive amounts of unstructured data in near real-time to discern critical insights that help clients make smarter, more informed decisions. We combine the best of finance, data, and technology to create a truly unique value proposition for hedge funds, asset managers, financial services firms, and major corporations.The Opportunity M Science is seeking a motivated and curious Data Scientist Intern to join our Summer 2027 Internship Program. This internship is designed for students passionate about alternative data, AI, and data science/engineering. You will gain hands-on experience building and testing data pipelines, developing AI/ML tools, and applying statistical models to real-world alternative datasets. Interns at M Science work alongside experienced data scientists and engineers to support impactful research used by hedge funds, asset managers, and other institutional clients.Key ResponsibilitiesAssist in building and optimizing data ingestion pipelines using Python, SQL, Databricks, and Spark.Support the development of AI/ML workflows and agents for automated insights, data analysis, and forecasting.Contribute to data testing, validation, and quality assurance, including anomaly detection and error handling.Process, cleanse, and verify the integrity of large and complex datasets.Help evaluate new alternative data assets for potential research and investment use.Work with senior data scientists to implement and document unit-tested analytics tools and models.Collaborate with cross-functional teams to support analytics, research, and client deliverables.What You’ll LearnHow to leverage alternative data and AI tools in investment research and analytics.Best practices in data engineering, testing, and pipeline optimization.Application of Python, SQL, PySpark, and Databricks for large-scale data processing.Fundamentals of statistical modeling and machine learning, including multivariate and ensemble methods.Introduction to AI/LLM frameworks and agent-based workflows.Preferred QualificationsCurrently pursuing a Bachelor’s or Master’s degree in Computer Science, Data Science, Statistics, Mathematics, Engineering, or a related quantitative field anticipated to graduate between December 2027 and May 2028Strong programming skills in Python and SQL; experience with PySpark/Spark is a plus.Familiarity with cloud platforms (AWS, Databricks) and distributed data processing.Interest or exposure to machine learning, AI agents, or large language models.Highly analytical, detail-oriented, and able to work independently and collaboratively.Strong written and verbal communication skills.Nice-to-Have SkillsExperience with LLM orchestration frameworks like LangChain or LangGraph.Familiarity with named entity resolution or other NLP methods.Prior experience in data testing, anomaly detection, or production-quality code.Why M Science?Hands-on experience with real alternative datasets and AI tools used in financial research.Work alongside senior data scientists and engineers on impactful projects.Hybrid work model offering both in-office collaboration and flexibility.Exposure to cutting-edge analytics in finance, technology, and data science.Hourly Rate: Up to $45/hr USD