أبلاي إيدج ابدأ البحث عن عمل

Network CCL Engineer

Gateworth Group · Dubai, United Arab Emirates

قدّم وتابع مع أبلاي إيدج
Position: Network CCL Engineer Compensation: Competitive Salary + Bonus + BenefitsLocation: Dubai, United Arab Emirates OverviewGateworth Group is supporting a high‑growth technology business that is building advanced accelerator hardware and large‑scale AI infrastructure, and they’re looking for a Network CCL Engineer who can operate deep in the performance layer. This role focuses on collective communication, accelerator runtime, heterogeneous device communication and distributed AI workloads. It suits someone who enjoys low‑level engineering, understands how accelerators behave at scale, and can design communication paths that perform across multi‑chip and multi‑rack environments. You’ll work closely with hardware, runtime and distributed systems teams to build communication foundations for next‑generation AI platforms.Main Responsibilities• Develop collective communication primitives for accelerator‑based systems.• Build communication support for multi‑device and rack‑level inference environments.• Design heterogeneous communication paths across accelerator and GPU devices.• Create the runtime components for device discovery, rank management, topology awareness, queues, streams, events and synchronization.• Integrate transport mechanisms across PCIe, RDMA‑based networking, custom fabrics and device‑to‑device communication.• Collaborate with hardware, runtime, compiler and distributed systems teams on scalable communication solutions.• Diagnose and optimise communication bottlenecks including bandwidth, latency, congestion and synchronization overhead.Qualifications • 5+ years in distributed systems, runtime engineering, networking or accelerator‑level development.• Strong C/C++ and Python programming experience.• Solid understanding of collective communication operations (AllReduce, ReduceScatter, AllGather, AllToAll, Broadcast, Send/Recv, Barrier).• Familiarity with libraries such as NCCL, RCCL, oneCCL, MPI, UCX or libfabric.• A good understanding of GPU/NPU memory behaviour, memory registration, peer‑to‑peer transfers and heterogeneous device communication.• Experience with runtime concepts including queues, streams, events, command submission, synchronization and memory management.• Understanding of hardware‑software interaction: DMA engines, device memory, interrupts, doorbells, fences and completion signalling.• Knowledge of high‑performance networking technologies (InfiniBand, RoCE).• Experience with distributed training frameworks or custom communication backends is a bonus.• Exposure to communication optimisation for LLM workloads (tensor parallelism, MoE, KV‑cache movement) would be advantageous.• Experience integrating communication layers with modern distributed inference frameworks would be helpful. ApplyApplicants meeting this criterion and looking for a progressive and challenging opportunity should submit an application via the apply link. If you have any further questions, you can reach out to apply@gateworth.com, quoting the GWG reference - #8356