ML/MLOps Engineer
Quantum Talent Group · Abu Dhabi Emirate, United Arab Emirates
Apply & track with Apply EdgeSenior Machine Learning / MLOps EngineerLocation: Abu Dhabi, UAE | On-siteAbout the opportunityWe are hiring for a rapidly growing advanced AI organisation based in Abu Dhabi, building and deploying sophisticated machine learning and generative AI systems for complex, high-impact real-world applications.The team operates at the intersection of applied AI research, machine learning engineering and production infrastructure, taking models from experimentation through to scalable, secure production environments.We are looking for a Senior Machine Learning / MLOps Engineer who can bridge the gap between model development and production. Depending on your background, you may lean more heavily toward ML engineering, GenAI or ML infrastructure, but you should be comfortable owning significant parts of the end-to-end machine learning lifecycle.What you'll work onDesign, develop and productionise advanced machine learning and deep learning systemsBuild and deploy LLM and NLP pipelines, including fine-tuning, semantic search and Retrieval-Augmented Generation (RAG)Deploy and scale large language models using modern inference frameworks such as vLLM, Triton or TGIDevelop machine learning solutions across areas such as forecasting, classification, anomaly detection, risk modelling and optimisationBuild automated pipelines for training, fine-tuning, evaluation, versioning, deployment and retrainingContainerise and orchestrate ML workloads using Docker and KubernetesImplement reproducible ML workflows using platforms such as MLflow, Kubeflow or equivalent toolingDesign data processing pipelines covering both structured and unstructured dataImprove inference performance through techniques such as quantisation, distillation, pruning and distributed or multi-GPU inferenceMonitor production models for performance, drift, latency, throughput and resource utilisationBuild reliable CI/CD and infrastructure automation for machine learning systemsWork closely with researchers, data scientists, software engineers and domain experts to transform prototypes into robust production solutionsDevelop systems that operate within environments where security, reliability, explainability and engineering rigour are criticalWhat we're looking for5+ years of experience in Machine Learning Engineering, MLOps, ML Infrastructure or a closely related fieldStrong track record of deploying machine learning models into productionExpert-level proficiency in PythonStrong knowledge of modern ML/deep learning frameworks such as PyTorch, TensorFlow and Scikit-learnExperience with transformer-based architectures and modern NLP/GenAI ecosystems such as Hugging Face, Llama-family models, GPT-style models or equivalentHands-on experience with Docker, Kubernetes and production ML orchestrationExperience with MLflow, Kubeflow, SageMaker Pipelines or comparable MLOps platformsStrong understanding of model serving, inference optimisation, monitoring and lifecycle managementExperience building scalable data and model pipelinesStrong understanding of software engineering and algorithmic fundamentalsAbility to work across ambiguous technical problems and quickly adapt ML techniques to specialised domainsAny of the following would be advantageous:Production deployment of large language modelsRAG, semantic search, embeddings or document intelligenceFine-tuning and optimisation of transformer modelsTime-series forecasting, anomaly detection or predictive modellingAWS machine learning infrastructure, including SageMaker, EC2 or EKSOn-premise or highly secure ML deployment environmentsDistributed training or inference using tools such as DeepSpeed, FSDP or AccelerateInfrastructure as Code and ML-focused CI/CDC/C++ experience for performance-sensitive systemsWhy consider this role?This is an opportunity to work on technically challenging AI systems where models move beyond prototypes and are deployed against meaningful, real-world problems.You'll have access to strong engineering talent, modern AI infrastructure and substantial compute resources while working across LLMs, traditional machine learning and large-scale production AI systems.