أبلاي إيدج ابدأ البحث عن عمل

MLOps Engineer

Inception42 · Abu Dhabi Emirate, United Arab Emirates

قدّم وتابع مع أبلاي إيدج
MLOps EngineerLocation: Abu Dhabi, UAE Employment type: Full-timeInception42, a G42 company, is the region’s leading innovator of AI-powered domain-specific as well as industry-agnostic products, built on a rich heritage of research and development. Within the G42 ecosystem, Inception42 functions as the core intelligence layer – transforming data and compute infrastructure into real-world, applied AI solutions. Beyond its commercial endeavors, Inception42 is committed to creating positive societal impact. For more information, please visit www.inceptionai.aiOverviewWe are looking for an MLOps Engineer to build and operate the systems that move machine learning from development into reliable production use. You will work across model training, validation, packaging, deployment, observability, and infrastructure automation, creating repeatable workflows that improve delivery speed without compromising security or control. The role works closely with data scientists, ML engineers, software engineers, architects, product teams, and platform teams in an environment where production quality, scale, and practical deployment matter.What You’ll OwnDesign, build, and operate end-to-end ML pipelines covering data preparation, training, validation, packaging, release, deployment, and production monitoring.Build CI/CD workflows for machine learning systems, with automated testing, quality gates, environment promotion, release controls, rollback, and reproducible builds.Own model lifecycle controls across experiment tracking, model and dataset versioning, registries, approvals, deployment governance, lineage, and production-readiness checks.Deploy and operate model-serving and inference workloads across production and non-production environments using containerised, cloud-native platforms.Create reusable platform components, templates, standards, and paved paths that reduce manual work and help teams move models from research to production consistently.Instrument ML systems for model quality, data and concept drift, latency, throughput, availability, infrastructure health, logs, traces, alerts, and service-level performance.Automate infrastructure and environment provisioning through infrastructure as code, configuration management, and disciplined controls for access, secrets, networking, storage, and compute.Debug and optimise pipelines, model-serving paths, and platform services at component and system level to improve reliability, scalability, performance, and cost efficiency.Embed security, compliance, access control, auditability, and data-governance requirements into ML platforms, automated workflows, and deployment processes.Collaborate closely with data science, engineering, product, and platform teams to resolve production issues, improve developer experience, document operating practices, and evolve the MLOps capability.What We’re Looking ForStrong engineering fundamentals across software delivery, machine learning systems, cloud infrastructure, automation, reliability, and production operations.A track record of building and operating production ML or data platforms, with clear ownership of deployment quality, reliability, security, and operational outcomes.Hands-on experience building CI/CD pipelines and automation using Python and shell scripting or comparable languages and tooling.Practical experience with containers and orchestration, including Docker and Kubernetes, and the ability to diagnose issues across applications, clusters, networks, storage, and compute.Working knowledge of Azure services for ML and application workloads, such as Azure Machine Learning, Azure AI Foundry, AKS, Azure DevOps, storage, networking, and managed data services, or equivalent cloud platforms.Experience with ML lifecycle and workflow capabilities such as experiment tracking, model registries, feature and dataset versioning, orchestration, and model serving using MLflow, Kubeflow, Airflow, DVC, or equivalent tools.Strong observability and debugging skills across model behaviour, data quality, drift, inference performance, distributed services, and infrastructure health.Clear communication and product judgement, with the ability to work through ambiguity, make pragmatic trade-offs, and collaborate effectively across research, engineering, platform, and product teams.Nice to HaveExperience deploying generative AI or LLM workloads, including prompt and model versioning, evaluation, guardrails, GPU environments, or high-performance inference.Exposure to distributed data platforms, search systems, SQL or NoSQL databases, time-series storage, feature stores, or large-scale training environments.Experience with Terraform or Bicep and observability platforms such as Prometheus, Grafana, ELK, Datadog, or Langfuse.