Senior LLMOps / AI Platform Engineer
Hire Rightt - Executive Search & HR Advisory · Dubai, United Arab Emirates
Apply & track with Apply EdgePosition: Senior LLMOps / AI Platform Engineer
The role combines LLM inference, GPU optimization, Kubernetes, cloud infrastructure, observability, RAG, and AI platform engineering.
Responsibilities
Deploy and operate self-hosted LLMs using vLLM, SGLang, Ollama.Optimize LLM inference for latency, throughput, concurrency, GPU memory, KV cache, and cost.Manage GPU workloads across multiple NVIDIA GPUs.Deploy and maintain AI services on Kubernetes / AWS EKS using Docker and Helm.Implement LLM reliability mechanisms including health checks, monitoring, automated recovery, and model restart/refresh strategies.Implement observability using Langfuse/LangSmith, OpenTelemetry, Prometheus, and Grafana.Deploy and optimize RAG systems, embedding models, and vector databases such as Qdrant, Milvus.Support AI agents and workflows built with LangChain and LangGraph.Build and maintain CI/CD pipelines for AI services and infrastructure.Troubleshoot production issues across LLMs, GPUs, Kubernetes, networking, and AI applications.
Requirements
Strong Python and FastAPI experiencevLLM, Hugging Face and self-hosted LLM deploymentKubernetes, Docker, Helm and AWSNVIDIA GPU inference and performance optimizationLangChain / LangGraph/LangSmithRAG, embeddings and vector databases (Qdrant)LLM observability and monitoringPostgreSQL / Redis GitHub Actions / CI/CDStrong production troubleshooting skillsEmail CVs to: mahin@hirerightt.com