Apply Edge Start your job search

AI Infrastructure Hardware Support Engineer-Onsite

Safeguard Global Recruiting · Jakarta, Indonesia

Apply & track with Apply Edge

Role Purpose:Provide hands-on hardware troubleshooting and support for NVIDIA-based AI/HPC infrastructure, including A100, H100, B200, B300, GB200 and GB300 platforms.

Key Responsibilities

Troubleshoot GPU servers, AI clusters, GPUs, CPUs, memory, system boards, NVLink/NVSwitch, NICs, DPUs, storage, power and cooling.Analyse BMC logs, NVIDIA Xid/ECC errors and hardware telemetry using NVIDIA SMI, DCGM, IPMI and OEM tools.Troubleshoot InfiniBand, Ethernet, RoCE, cables and transceivers.Perform hardware health checks, component replacement, testing and post-repair validation.Resolve BIOS, firmware, BMC, driver and hardware compatibility issues.Coordinate hardware replacements and escalations with OEMs/NVIDIA.Maintain incident records, troubleshooting guides and RCA reports.Support shift-based/on-call requirements when needed.

Required Skills

3+ years of Data Centre, Server or HPC hardware support experience.Strong enterprise server and component-level troubleshooting skills.Knowledge of NVIDIA GPUs and multi-GPU systems.Working knowledge of Linux and server-management tools.Understanding of networking, storage, rack power and cooling.GPU/AI infrastructure experience is highly preferred.Strong troubleshooting, fault isolation and incident-management skills.