Apply Edge Start your job search

Senior AI/ML Engineer ( US Helthcare )

Ekshvaku Tech Innovations · Hyderabad, Telangana, India

Apply & track with Apply Edge

Senior AI/ML Engineer — Clinical AI Quality & LLMOps - HealthcareRole PurposeWe are looking for a Senior AI/ML Engineer — Clinical AI Quality & LLMOps to establish dedicated ownership of the quality, reliability, evaluation, optimization, and operational maturity of DoctusMind's clinical AI capabilities.This role will sit at the intersection of AI/ML Engineering, Clinical AI Quality, LLMOps, AI Evaluation, Prompt Engineering, and Production AI Operations.The person will work closely with Product, Engineering, QA, Clinical, and other cross-functional teams to ensure that DoctusMind's AI systems are clinically reliable, measurable, observable, cost-efficient, and production-ready.The role is not limited to prompt engineering. The expectation is that this person will build the systems, processes, evaluation frameworks, and engineering practices required to operate clinical AI reliably at scale.The initial focus will include AI stabilisation, clinical AI quality, Chronic Care, and longitudinal intelligence, while establishing the foundation for future specialty, multimodal, agentic, retrieval, memory, and other clinical AI capabilities.What Success Looks LikeThe person will be accountable for five executive-level outcomes during the first 3–6 months.AI Quality & StabilizationProduction AI Release GovernanceAI Observability & Operational PerformanceClinical AI Capability Production ReadinessCore Responsibilities1. Prompt Engineering & ManagementDesign, optimize, and maintain production-grade prompts.Establish prompt versioning and lifecycle management.Build prompt regression suites.Analyze prompt-related failures.Optimize prompts for quality, consistency, latency, and cost.Maintain traceability between prompt changes and evaluation outcomes.2. AI Evaluation & Clinical QualityBuild AI evaluation frameworks for accuracy, relevance, consistency, safety, and reliability.Develop golden datasets and clinical test cases.Create automated evaluation workflows.Establish severity classification for AI failures.Perform systematic failure analysis and root-cause analysis.Monitor clinical AI quality continuously.Partner with Clinical and QA teams to validate high-risk workflows.3. AI Regression & Release ManagementEstablish AI regression testing as part of the release lifecycle.Define evaluation gates for prompts, models, retrieval, and configurations.Ensure AI changes cannot reach production without appropriate validation.Maintain AI version traceability.Establish rollback and post-release validation processes.4. LLMOps & AI ObservabilityBuild and maintain AI-specific monitoring.Track quality, latency, failures, token consumption, and cost.Develop dashboards and operational reporting.Establish alerts for significant deviations.Support AI production incident investigation.Identify AI quality and performance drift.5. AI Experimentation & A/B TestingMaintain environments for prompt and model experimentation.Conduct controlled A/B tests.Compare prompts, models, retrieval strategies, and AI workflows.Define experiment methodology and success criteria.Document findings and recommendations before production adoption.6. Model Benchmarking & SelectionBenchmark existing and emerging models against DoctusMind use cases.Evaluate models based on:Clinical qualityReliabilitySafetyLatencyCostScalabilityRecommend model selection and routing strategies.Evaluate fallback strategies for production resilience.7. Cost & Token OptimizationAnalyze LLM/API consumption.Optimize prompt and context length.Evaluate caching opportunities.Optimize model selection and routing.Identify unnecessary or inefficient AI calls.Establish ongoing AI cost-performance monitoring.8. Context Engineering, Retrieval & MemoryOptimize how patient information, conversations, care plans, and clinical context are supplied to AI systems.Improve retrieval of relevant patient history and clinical context.Support RAG, embeddings, vector databases, and retrieval pipelines.Contribute to cross-session memory and longitudinal intelligence.Evaluate context quality and relevance as part of AI evaluation.9. Clinical AI Capability ProductionizationEvaluate new AI capabilities before production.Support Chronic Care and longitudinal intelligence initiatives.Contribute to specialty-specific AI capabilities.Support multimodal clinical AI capabilities.Evaluate AI agents, tools, reasoning capabilities, and emerging technologies.Establish production-readiness criteria for new capabilities.10. Continuous AI ImprovementEstablish feedback loops using clinician, patient, QA, and system signals.Analyze production feedback and recurring failure patterns.Convert production learnings into evaluation datasets and regression tests.Continuously improve AI quality and operational performance.Candidate ProfileRequired Experience5+ years of software engineering / ML engineering experience, including:2–3+ years building and operating production LLM/Generative AI systemsHands-on experience with AI evaluation and regression testingProduction LLMOps experienceAI observability and monitoringModel benchmarking and experimentationPrompt engineeringProduction AI troubleshootingAPI-based AI/LLM applicationsPythonAI performance and cost optimizationHealthcare Experience — RequiredCandidates should have demonstrated experience in one or more of:Healthcare AIClinical AIHealthcare technologyClinical NLPHealthcare data platformsClinical decision-support systemsHealthcare-focused LLM applicationsHealthcare/clinical AI experience should be treated as a core screening criterion, not simply a preferred qualification.Technical SkillsStrong hands-on experience with several of the following:LLM / GenAIOpenAI / Gemini / Claude or equivalent LLM platformsPrompt engineeringStructured generationFunction/tool callingAI agentsModel evaluationAI EvaluationGolden datasetsAutomated evaluationLLM-as-a-judge approachesRegression testingEvaluation pipelinesError classificationQuality benchmarkingLLMOps / MLOpsAI observabilityModel/version managementProduction monitoringCI/CDRelease governanceExperiment trackingRAG / ContextRAGEmbeddingsVector databasesRetrieval optimizationContext engineeringMemory architecturesEngineeringPythonREST/API integrationsCloud platformsLogging/monitoringData pipelinesAutomated testingOptimizationToken optimizationPrompt optimizationModel routingCachingLatency optimizationAI cost optimioptimisationzationProduction LLM/GenAI experience, clinical/healthcare AI experience, AI evaluation and regression, LLMOps/observability, and AI quality and optimisation.