Senior Machine Learning Engineer
Inception42 · Abu Dhabi Emirate, United Arab Emirates
قدّم وتابع مع أبلاي إيدجSenior Machine Learning EngineerLocation: Abu Dhabi, UAE Inception42, a G42 company, is the region’s leading innovator of AI-powered domain-specific as well as industry-agnostic products, built on a rich heritage of research and development. Within the G42 ecosystem, Inception42 functions as the core intelligence layer – transforming data and compute infrastructure into real-world, applied AI solutions. Beyond its commercial endeavors, Inception42 is committed to creating positive societal impact. For more information, please visit www.inceptionai.aiOverviewWe are looking for a Senior Machine Learning Engineer to turn advanced machine learning work into reliable, scalable products. You will own the engineering path from data and experimentation through training, evaluation, deployment, and production monitoring, working closely with applied scientists, data scientists, MLOps engineers, software engineers, and product teams. The environment is technically demanding and deployment-focused: models must perform under real-world constraints for latency, reliability, security, governance, and cost.What You’ll OwnDesign, build, and scale production machine learning systems across data preparation, feature engineering, training, evaluation, deployment, and monitoring.Translate research prototypes into maintainable product capabilities, defining production-readiness criteria and closing gaps in data quality, performance, reliability, and operability.Build reusable training and evaluation pipelines with disciplined experiment tracking, model and dataset versioning, validation, benchmarking, and reproducibility.Engineer batch and online inference paths, model-serving components, APIs, and feature workflows that meet product requirements for latency, throughput, availability, and cost.Optimise models and systems across accuracy, robustness, memory, compute efficiency, inference speed, and scalability using evidence from profiling and real workloads.Instrument production ML systems for model quality, drift, data integrity, latency, throughput, availability, and operational health; diagnose issues across the full stack.Collaborate closely with MLOps, platform, and software engineering teams on CI/CD, containerisation, orchestration, release controls, rollback, and production operations.Embed security, privacy, governance, auditability, and compliance requirements into data, model, and deployment workflows from the start.Drive architecture and technical decisions for ML components, balancing product needs with maintainability, reliability, performance, and delivery speed.Raise engineering quality through design reviews, code reviews, technical documentation, reusable standards, and practical mentorship for other engineers.What We’re Looking ForStrong software engineering fundamentals and fluency in Python, with a track record of building tested, maintainable systems rather than isolated notebooks or prototypes.Experience shipping and operating production machine learning systems, including ownership of deployment quality, reliability, observability, and operational outcomes.Strong understanding of machine learning fundamentals and practical depth in one or more areas such as deep learning, natural language processing, computer vision, recommendation, or generative AI.Experience with modern ML frameworks such as PyTorch, TensorFlow, or scikit-learn, and sound judgment about selecting tools for the problem rather than applying a fixed stack.Working knowledge of data and training workflows, including SQL, data validation, feature engineering, distributed processing, and efficient use of large datasets.Practical experience with model serving, APIs, containers, orchestration, CI/CD, and cloud infrastructure in production environments.Strong debugging and systems-thinking skills across data, model behaviour, application code, distributed services, and infrastructure.Clear communication and product judgment, with the ability to work through ambiguity, make pragmatic trade-offs, and collaborate across research, engineering, platform, and product teams.Nice to HaveExperience with large language models, retrieval-augmented generation, agentic systems, evaluation frameworks, guardrails, or high-performance inference.Exposure to Azure Machine Learning, Azure AI Foundry, Databricks, AKS, or equivalent cloud-native AI and data platforms.Experience with distributed training, GPU optimisation, feature stores, vector search, or large-scale model-serving infrastructure.Contributions to reusable ML platforms, open-source systems, applied research, or technical standards that have improved engineering practice beyond a single project.