أبلاي إيدج ابدأ البحث عن عمل

Inference Server Engineer

Gateworth Group · Dubai, United Arab Emirates

قدّم وتابع مع أبلاي إيدج
Position: Inference Server EngineerCompensation: Competitive salary plus family benefits & variableLocation: Dubai, United Arab EmiratesOverviewGateworth Group is assisting a fast‑growing global organisation in building advanced AI acceleration platforms designed for large‑scale inference. They’re developing custom hardware and software that power next‑generation datacentre systems, and they’re looking for an engineer who understands LLM inference servers at a deep, internal level.This role sits at the intersection of model execution, runtime optimisation, distributed inference, and backend integration. You’ll be responsible for enabling their accelerator to function as a first‑class backend inside modern serving frameworks, ensuring high‑performance execution across dense and MoE workloads.Main Responsibilities• Build backend support for the company’s accelerator inside leading LLM inference servers (e.g., vLLM, SGLang, TensorRT‑LLM).• Implement device registration, runtime execution paths, memory handling, custom operators, and KV‑cache integration.• Enable both dense and MoE inference, including attention execution, expert routing, expert scheduling, and multi‑device distributed execution.• Profile, debug, and optimise inference performance across NPU and mixed NPU–GPU environments, including correctness, performance, and stress testing.Qualifications • 5+ years of relevant experience, with a background in Computer Science, Computer Engineering, or equivalent industry experience.• Strong programming skills in C/C++ and Python.• Hands‑on experience modifying or extending LLM inference server internals (not just deploying them).• Experience running and scaling workloads on heterogeneous clusters (CPU + GPU) using distributed inference or training strategies.• Familiarity with PyTorch internals, custom operator registration, tensor execution, model loading, and integration with Hugging Face models.• Understanding of distributed inference techniques such as tensor parallelism, expert parallelism, data parallelism, multi‑device execution, and multi‑node serving.Bonus Qualifications:• Prior contributions or merged PRs to vLLM, SGLang, TensorRT‑LLM, or similar projects.• Knowledge of accelerator runtime and compiler stacks.• Experience with GPU/NPU memory models, allocation, movement, and heterogeneous device execution.• Familiarity with distributed communication libraries (NCCL, RCCL, oneCCL, MPI, UCX, libfabric).• Experience with CUDA or Triton kernel development.ApplyApplicants meeting this criterion and looking for a progressive and challenging opportunity should submit an application via the apply link. If you have any further questions, you can reach out to apply@gateworth.com, quoting the GWG reference - #8355