Senior AI Engineer
Gateworth Group · Dubai, United Arab Emirates
Apply & track with Apply EdgePosition: Senior AI Engineer Compensation: Competitive salary plus family benefits & variable Location: Dubai, United Arab Emirates OverviewGateworth Group is partnering with a technology organisation in the UAE that is scaling its AI engineering function and building a specialist team focused on high‑performance model optimisation. They’re investing heavily in advanced AI infrastructure and are seeking senior engineers who can work deep inside modern LLMs, improve inference behaviour, and shape how large‑scale models run in production.This role suits someone who enjoys complex, hands‑on engineering, understands how transformer architectures behave at scale, and can move confidently between model internals, systems performance, and deployment‑level optimisation. You’ll work across analysis, tuning, benchmarking and architectural decision‑making, contributing to next‑generation AI systems.Main Responsibilities• Improve and optimise LLM inference performance across distributed, multi‑chip and multi‑node environments• Apply strong understanding of transformer architectures, including dense and Mixture‑of‑Experts (MoE) models• Benchmark leading LLMs (LLaMA, Mistral, Qwen, DeepSeek) across varied hardware stacks• Design and implement attention‑level optimisations (Flash Attention, grouped‑query, sliding‑window)• Deliver model‑level optimisation including quantisation (INT8/FP8), KV‑cache strategies, batching and parallelism• Work closely with hardware, systems and compiler teams to co‑design efficient inference pipelines• Build and maintain benchmarking frameworks to measure latency, throughput and scaling behaviour• Evaluate architectural trade‑offs and contribute to deployment strategies for large‑scale environments• Stay current with research across LLM architectures, inference optimisation and performance engineeringQualifications• Strong understanding of transformer architectures, LLM internals and both dense/MoE models• Hands‑on experience with modern LLMs (LLaMA, Mistral, Qwen, DeepSeek) and attention‑level optimisation• Practical experience with inference optimisation: quantisation (INT8/FP8), KV‑cache strategies, batching, pruning and parallelism• Strong Python skills with PyTorch or JAX, plus experience profiling and debugging performance bottlenecks• Background in distributed systems, large‑scale inference workloads and system‑level optimisation across hardware and runtime layers• Ideally 8+ years in deep learning, AI systems or performance engineering, with exposure to datacenter‑scale inference (e.g., vLLM) and hardware‑aware optimisationApplyCandidates who meet the above criteria and are seeking a progressive, technically challenging role are invited to apply via the link provided. For further questions, contact apply@gateworth.com quoting reference #8341.