أبلاي إيدج ابدأ البحث عن عمل

Senior AI Engineer

Gateworth Group · Dubai, United Arab Emirates

قدّم وتابع مع أبلاي إيدج
Position: Senior AI Engineer – LLM SystemsCompensation: Competitive salary plus family benefits & variableLocation: Dubai OverviewGateworth Group is supporting a technology organisation in the UAE that is expanding its AI engineering capability and building out a specialised team focused on large‑scale model performance. The company is investing heavily in advanced AI systems and is hiring senior engineers who can work deep in the internals of modern LLMs, optimise inference behaviour, and contribute to high‑performance model deployment across complex environments.This role suits someone who enjoys highly technical work, understands how transformer‑based models behave at scale, and can move comfortably between research, engineering, and system‑level optimisation. You’ll work across model analysis, performance tuning, benchmarking, and architecture decisions, helping shape next‑generation AI infrastructure.Main Responsibilities• Analyse, profile, and optimise LLM inference performance across distributed, multi‑chip or multi‑node systems• Apply deep understanding of transformer architectures, including dense and Mixture‑of‑Experts (MoE) models• Evaluate and benchmark leading LLMs (LLaMA, Mistral, Qwen, DeepSeek) across different hardware environments• Design and implement optimisations for attention mechanisms such as Flash Attention, grouped‑query attention, and sliding‑window attention• Work on model‑level optimisation techniques including quantisation (INT8/FP8), KV‑cache management, batching, and parallelism strategies• Collaborate with hardware, systems, and compiler teams to co‑design efficient inference pipelines• Build and maintain benchmarking frameworks to evaluate latency, throughput, and scaling behaviour• Analyse trade‑offs between model architecture choices and system‑level performance, contributing to deployment strategies for large‑scale environments• Stay current with research in LLM architectures and inference optimisationQualifications• Strong understanding of transformer architectures and LLM internals• Hands‑on experience with multiple modern LLMs (LLaMA, Mistral, Qwen, DeepSeek)• Deep knowledge of dense and Mixture‑of‑Experts (MoE) architectures• Familiarity with attention mechanisms and their optimisation strategies• Experience with inference optimisation techniques (quantisation, pruning, KV‑caching, batching)• Strong Python skills and experience with ML frameworks such as PyTorch or JAX• Experience with distributed systems and large‑scale inference workloads• Ability to profile and debug performance bottlenecks across hardware and software stacks• Strong systems thinking with the ability to work across model, runtime, and hardware layersPreferred:• 8+ years’ experience in deep learning, AI systems, or performance engineering• Experience working close to hardware• Familiarity with parallelism strategies (tensor, pipeline, expert parallelism)• Experience with datacenter‑scale deployment and inference servers (e.g., vLLM)• Background in performance engineering or systems optimisationApplyApplicants meeting this criterion and looking for a progressive and challenging opportunity should submit an application via the apply link. If you have any further questions, you can reach out to apply@gateworth.com, quoting the GWG reference - #8341