Hardware Engineer
Yotta Data Services Private Limited · Navi Mumbai, Maharashtra, India
قدّم وتابع مع أبلاي إيدجJob Description – Hardware EngineerExperience: 5+ YearsRole Overview:We are looking for an experienced and hands-on Hardware Engineer to support the operations, maintenance, troubleshooting and lifecycle management of server hardware across our infrastructure. The role will involve working with both AI/GPU-based servers and conventional CPU-based enterprise servers in mission-critical data center environments.The ideal candidate should have strong experience in server hardware break-fix, fault diagnosis, component replacement, spare-parts management, RMA processes and coordination with OEM/ODM support teams.Key Responsibilities:1. Server Hardware Operations & MaintenancePerform end-to-end troubleshooting, maintenance and repair of AI/GPU and conventional enterprise servers.Ensure high availability and reliability of server hardware deployed across the infrastructure.Monitor and identify hardware faults, degradation and recurring failure patterns.Perform hardware diagnosis and identify faulty components for replacement.Ensure servers are tested and operational after repair, replacement or maintenance activities.Work within defined SLAs to ensure timely resolution of hardware-related incidents.2. Server Hardware Break-FixHandle the complete hardware break-fix lifecycle, including:Fault detection and diagnosisFaulty component identificationSpare allocationOnsite component replacementServer testing and service restorationFaulty part removal and taggingRMA coordination and closureTroubleshoot and replace server Field Replaceable Units (FRUs), including:GPUsCPUsMotherboardsMemoryNICsPower Supply Units (PSUs)FansStorage and other server components3. AI/GPU and Enterprise Server SupportProvide hardware support for NVIDIA GPU-based servers and compute nodes.Work on high-density AI and HPC server environments.Support GPU server platforms, including HGX/DGX architecture and other GPU-based server platforms.Perform troubleshooting and replacement of GPU cards and associated server components.Support conventional enterprise and CPU-based servers, including:General-purpose compute serversApplication and database serversVirtualization serversHigh-performance compute servers4. Spare Parts & Inventory ManagementMaintain and manage server hardware spares required for break-fix activities.Ensure proper receipt, inspection, storage and issue of server components.Maintain accurate records of spare inventory and component movement.Monitor availability of critical server FRUs and escalate requirements for replenishment.Support inventory reconciliation of replaced, repaired and available spare components.Ensure critical AI/GPU server components are available as per operational requirements.5. Faulty Part & RMA ManagementIdentify, tag and maintain proper records of faulty or replaced components.Follow the defined process for segregation and storage of faulty hardware.Coordinate with OEMs/ODMs for raising and tracking RMA cases.Prepare faulty components for shipment to designated OEM/ODM service centers.Track repair and replacement status of faulty components.Ensure repaired or replacement components are received and appropriately updated in the inventory.Maintain accurate documentation to avoid unaccounted or misplaced components.6. OEM/ODM & Vendor CoordinationCoordinate with OEMs, ODMs and hardware service partners for technical support and issue resolution.Follow up on delayed parts, replacement requests and unresolved hardware issues.Support warranty and service-related activities.Escalate critical or recurring hardware failures to the relevant internal and external teams.Coordinate with central engineering teams and onsite support teams for timely issue resolution.7. Incident Management & DocumentationRespond to server hardware incidents within defined response and resolution timelines.Maintain detailed records of hardware faults, repairs, replacements and RMA activities.Document troubleshooting steps and resolutions for recurring issues.Follow established operational processes, escalation procedures and SLAs.Participate in shift or standby support for critical 24x7 environments, as required.Required Skills & Experience:5+ years of experience in server hardware operations, data center hardware support or enterprise server infrastructure.Strong hands-on experience in server hardware troubleshooting and break-fix operations.Strong understanding of server hardware architecture and components.Experience in diagnosing and replacing server components such as GPUs, CPUs, motherboards, memory, NICs, PSUs, fans and other FRUs.Experience with spare-parts management and hardware inventory.Experience in faulty part handling and RMA management.Experience coordinating with OEMs, ODMs and hardware service partners.Understanding of hardware incident management and SLA-driven support environments.Experience working in 24x7 mission-critical data center or enterprise infrastructure environments.Ability to troubleshoot issues independently and coordinate with multiple technical teams.Preferred Skills:Experience with one or more of the following will be an added advantage:NVIDIA GPU-based servers and compute infrastructureNVIDIA H100, H200, B200, B300, GB200 or GB300 platformsNVIDIA HGX/DGX architectureAI/HPC server environmentsHigh-density compute infrastructureSupermicroASUSGigabyteDellHPEGPUaaS, AI Cloud or large-scale AI data center environmentsEducational Qualification:Bachelor's degree or diploma in Computer Science, Electronics, Electrical Engineering, Information Technology, or a related technical discipline.Relevant hardware, server or OEM certifications will be an added advantage.