Apply Edge Start your job search

Manager, MLOPS, Machine Learning and Data Science

CareSet · United States

Apply & track with Apply Edge
About the CompanyCareSet is a thriving young healthcare data analytics company on a mission to unlock Medicare claims data for the pharmaceutical industry. We've built the analytical engine, now we need someone to build the financial engine to match it.About the RoleThis is a player-coach position for a senior technical leader who wants to do real work and build a team around it. You will own the end-to-end machine learning lifecycle on Medicare and Medicaid claims data: from raw line-level claims through feature engineering, model development, MLOps deployment on Databricks, and integration into analytical delivery pipelines that go directly to pharmaceutical and life sciences clients. You will also build and lead a team of machine learning engineers and biostatisticians, setting the technical standard for everything the team produces. The work demands both deep methodological expertise in real-world evidence and the engineering discipline to move models reliably from development into production.ResponsibilitiesMLOps and Model Lifecycle ManagementOwn the full model lifecycle on Databricks: data ingestion and validation, feature engineering, model training, evaluation, registration, deployment, monitoring, rollback, and retirement, with reproducibility and governance built into every stage.Build and maintain a versioned feature store to ensure reusable, consistently defined feature sets are available across models and use cases, reducing duplication and improving cross-model consistency.Establish model approval gates: define evaluation criteria and review standards a model must satisfy before advancing from experiment to registry to production; maintain documented evidence at each gate.Define and monitor model health indicators across accuracy, calibration, feature drift, data quality, and operational performance; set alert thresholds and own the response when models degrade in production.Own ML incident response: when a production model degrades or fails, lead the investigation, determine whether the response is rollback or retrain, execute the fix, and document the lesson.Establish release readiness standards: automated pipeline tests, reproducible builds, peer code review, and compliance documentation before any model ships to a client deliverable.Ensure all model development and deployment complies with CMS DUA obligations, CareSet's AI governance framework, and applicable data privacy and security standards.Machine Learning and Analytics DeliveryApply advanced machine learning and statistical methods to large-scale Medicare and Medicaid claims data to generate insights that inform pharmaceutical and life sciences decisions, including treatment effectiveness, patient journey mapping, disease progression, and real-world evidence generation.Design and build analytical frameworks grounded in sound epidemiological principles: appropriate cohort construction, longitudinal follow-up, bias control, and confounding adjustment on observational healthcare data.Perform large-scale feature engineering and data preparation using Python, PySpark, and SQL within Databricks, working across billions-row administrative health datasets with the query discipline and performance optimization that scale demands.Translate complex model outputs into transparent, clinically meaningful insights for pharmaceutical and life sciences stakeholders, using model interpretability and explainability methods appropriate to each use case and audience.Integrate diverse data sources including biomarker labs, biometric data, and CMS claims; reconcile privacy-linked data via tokenization into high-quality, AI-ready analytical assets; apply re-identification risk assessment and HIPAA safe harbor protocols across use cases including line of therapy, disease staging, persistency, and HCP targeting.Apply emerging AI capabilities where they add analytical value, including AI-assisted evidence synthesis, automated literature review, and intelligent querying of clinical knowledge, as complements to rigorous ML model development.Stay current with advances in machine learning, health informatics, and real-world evidence methodology, and bring relevant techniques into the team's practice as the work evolves.Team LeadershipBuild, mentor, and grow a team of machine learning engineers and biostatisticians: hiring, onboarding, technical goal-setting, peer code review culture, and performance management.Set and enforce team standards for code quality, peer review, model validation, and documentation; drive adoption of shared engineering practices and continuous quality improvement.Coordinate with delivery and engineering leads to align modeling work with client timelines and resource capacity.Partner with the CDAO on workforce planning, team structure, and technical governance as the data science function scales.Client and Cross-Functional CollaborationEngage directly with pharmaceutical and life sciences clients to frame analytical problems, design research plans, and translate model outputs into clear, defensible narratives for non-technical and executive audiences.Present complex findings and study design decisions to senior stakeholders; model the communication standard for the team.Support business development through technical scoping, methodology review, and feasibility analysis for new engagements.Serve as the internal authority on CMS data environments including the CCW and the VRDC, and ensure all work is conducted in full compliance with CMS DUA obligations.QualificationsMaster's or PhD in Biostatistics, Epidemiology, Statistics, Data Science, Computer Science, or a closely related quantitative field.10 or more years of professional experience in machine learning, data science, or applied biostatistics in healthcare or life sciences, with a track record of delivering production-grade models and evidence-generation work on large observational datasets.3 or more years of direct people management experience leading ML or data science teams, including hiring and performance management.Expert-level Python and PySpark proficiency, with specific depth in query optimization, handling data skew, and leveraging window functions on massive datasets.Deep, hands-on Databricks experience including Delta Lake, Unity Catalog, MLflow, and large-scale Spark-based data processing.Proven MLOps delivery at scale on Databricks: end-to-end model lifecycle management using PySpark MLlib or scalable packages such as SynapseML and XGBoost-on-Spark, model registry, deployment pipelines, drift monitoring, model approval workflows, and incident response.Demonstrated mastery of machine learning explainability frameworks to validate model outputs and communicate feature importance clearly to healthcare and life sciences stakeholders.Experience establishing responsible AI standards in a regulated environment: model governance, human oversight, auditability, and release readiness controls.Documented, hands-on experience working with CMS Chronic Condition Warehouse (CCW) data.Operational experience working within the VRDC secure research environment, including file handling, output review, and data governance compliance.Deep familiarity with Medicare Standard Analytic Files, Research Identifiable Files, and Medicaid T-MSIS data including their structure, limitations, and appropriate analytical use.Working knowledge of CMS DUA obligations and the governance constraints they impose on data use, model deployment, and client deliverables.