AI Infrastructure and Platform Operations Resident - strong exp in NVAIE, NIM, Triton endpoint support and troubleshooting ,Kubernetes administration and cluster operations,
Avensys Consulting · Riyadh, Saudi Arabia
Apply & track with Apply EdgeAvensys is a reputed global IT professional services company headquartered in Singapore. Our service spectrum includes enterprise solution consulting, business intelligence, business process automation, and managed services. Given our decade of success, we have evolved to become one of the top trusted providers in Singapore and service a client base across banking and financial services, insurance, information technology, healthcare, retail and supply chain.We are hiring for "AI Infrastructure and Platform Operations Resident - strong exp in NVAIE, NIM, Triton endpoint support and troubleshooting ,Kubernetes administration and cluster operations, GPU orchestration, PowerEdge, PowerScale F210/OneFS, AI networking, RoCE/RDMA, and storage performance" for Riyadh KSA Experience - 8 + YearsJob Summary:BCM and server lifecycle managementKubernetes administration and cluster operationsRun:ai resource pools, quotas, priorities, fair-share, reservations, and queue managementGPU health, utilization, capacity, and workload-placement monitoringPowerEdge, PowerScale F210/OneFS, AI networking, RoCE/RDMA, and storage performanceNVAIE, NIM, Triton endpoint support and troubleshootingLinux, incident management, change control, patching, upgrades, backup, and recoveryPerform recurring health reviews across servers, GPUs, management and infrastructure nodes, PowerScale, switching, storage paths, Kubernetes, BCM, Run:ai, and critical AI services.Maintain BCM provisioning, node baselines, configuration consistency, lifecycle coordination, and recovery procedures.Administer Kubernetes nodes, namespaces, workloads, GPU resources, and platform add-ons within the MOJ operating model.Operate H200 inference, H200 fine-tuning, and L40S RAG resource pools, including quotas, priorities, fair-share, reservations, fractional-GPU allocation, and approved preemption.Monitor utilization, queue time, idle capacity, throughput, contention, endpoint availability, and workload placement.Support PowerScale capacity and protection reviews and investigate AI-network, storage, GPU-fabric, and endpoint issues.Coordinate firmware, driver, operating-system, Kubernetes, BCM, Run:ai, and NVAIE maintenance planning with MOJ change processes.Provide first-line incident triage, impact assessment, escalation, recovery tracking, and post-incident improvement actions.Maintain runbooks for cluster health, provisioning, workload submission, GPU allocation, endpoint troubleshooting, patching, upgrades, incident response, and recovery.WHAT’S ON OFFERYou will be remunerated with an excellent base salary and entitled to attractive company benefits. Additionally, you will get the opportunity to enjoy a fun and collaborative work environment, alongside a strong career progression.To submit your application, please apply online or email your UPDATED CV in Microsoft Word format to saranya@aven-sys.com Your interest will be treated with strict confidentiality.CONSULTANT DETAILSConsultant Name: Saranya MAvensys Consulting Pte LtdEA Licence 12C5759Privacy Statement: Data collected will be used for recruitment purposes only. Personal data provided will be used strictly in accordance with the relevant data protection law and Avensys' personal information and privacy policy.