Apply Edge Start your job search

Senior HPC Engineer

Core42 · Abu Dhabi, Abu Dhabi Emirate, United Arab Emirates

Apply & track with Apply Edge
Senior HPC EngineerAbout UsCore42, a leader in AI-powered cloud and digital infrastructure, is driving transformative technology solutions globally. Leveraging advanced resources and partnerships, Core42 empowers clients to harness sovereign AI infrastructure, especially in sectors with stringent regulatory needs. With a mission to redefine digital transformation, we combine sovereign capabilities with scalable, high-performance compute infrastructure, positioning itself at the forefront of AI innovation in the Middle East and beyond.The opportunityWe are seeking a highly skilled Senior HPC Engineer to support the design, implementation, deployment, and ongoing operations of high-performance computing infrastructure. The role will be responsible for ensuring the availability, reliability, and performance of complex HPC environments spanning compute, storage, networking, InfiniBand, GPU resources, job scheduling, and supporting platforms.The ideal candidate will bring strong hands-on experience across Linux systems administration, HPC cluster deployment, network engineering, storage, workload management, and automation. You will work closely with internal engineering teams, customers, and technology vendors to troubleshoot complex issues, optimize infrastructure, and maintain highly available HPC environments supporting demanding workloads, including AI/ML and specialized industry applications.Your key responsibilitiesSupport the design, implementation, operation, and maintenance of Core42’s HPC infrastructure, including compute, storage, networking, InfiniBand, and associated management platforms.Install, configure, deploy, and maintain HPC clusters, including compute nodes, storage nodes, interconnects, and supporting infrastructure.Configure, manage, and troubleshoot InfiniBand and high-speed Ethernet networks, including NVIDIA Mellanox switches, subnet managers such as OpenSM, and routing configurations.Implement and maintain job scheduling and workload management platforms, with strong experience in LSF preferred and exposure to Slurm and/or PBS.Administer and support Red Hat Enterprise Linux and Windows Server environments within complex HPC and data center infrastructures.Support HPC storage environments, including IBM Spectrum Scale (GPFS), DDN GridScaler, NetApp SAN/NAS, and backup solutions such as IBM Spectrum Protect (TSM).Utilize HPC administration and provisioning toolkits, including xCAT, PXE, Kickstart, and other automated deployment technologies.Support and optimize GPU-enabled infrastructure, including NVIDIA H100 or newer architectures, and tune workloads for CPU and GPU resources.Monitor system, network, storage, and infrastructure health using tools such as Zabbix, Grafana, and other monitoring and observability platforms.Perform advanced troubleshooting and root cause analysis across complex HPC, network, storage, and operating system environments.Support virtualization platforms including VMware ESXi/vSAN, Citrix, VxRail, and KVM where required.Administer and support enterprise server platforms, including HPE ProLiant, Dell PowerEdge, and other compute infrastructure.Develop automation and operational tooling using Shell scripting and Python to improve efficiency, reliability, and repeatability.Create and maintain technical documentation, architecture diagrams, and operational procedures using tools such as Visio and Draw.io.Apply HPC security best practices, including encryption, firewalls, access controls, network security, and compliance requirements.Collaborate effectively with customers, vendors, and internal multidisciplinary teams to resolve technical issues and support infrastructure improvements.Support specialized HPC workloads and, where applicable, industry applications such as reservoir engineering simulators and petroleum software tools including tNavigator, Nexus-VIP, Eclipse, Intersect, IMPOWER, and PUMA.Participate in continuous improvement initiatives to enhance the performance, scalability, availability, and reliability of HPC environments.Qualifications:What we’re looking for(a) Required skills / qualificationsBachelor’s or Master’s degree in Computer Science, Engineering, Information Technology, or a related technical field.5+ years of experience in HPC operations, systems engineering, infrastructure engineering, network engineering, or a related technical role, preferably within an HPC environment.Strong technical background with extensive hands-on experience supporting multiple technologies and complex, highly available HPC environments.Strong experience with Red Hat Enterprise Linux systems administration and enterprise data center infrastructure.Hands-on experience deploying and administering HPC clusters, including compute, storage, networking, and high-speed interconnect components.Strong expertise in InfiniBand networking and Ethernet technologies, including the configuration and management of NVIDIA Mellanox InfiniBand switches.Experience with subnet managers such as OpenSM, routing configurations, network protocols, and troubleshooting methodologies within HPC environments.Experience with HPC job scheduling and workload management systems; LSF experience is strongly preferred, with Slurm and/or PBS experience also highly desirable.Familiarity with IBM Spectrum Scale (GPFS), IBM Spectrum Protect (TSM), DDN GridScaler, and NetApp SAN/NAS storage technologies.Experience with HPC provisioning and administration tools such as xCAT, PXE, and Kickstart.Knowledge of GPU infrastructure and workload optimization, with experience supporting NVIDIA H100 or newer GPU technologies considered desirable.Experience tuning workloads and infrastructure for specific CPU and GPU hardware configurations.Familiarity with VMware ESXi/vSAN, Citrix, VxRail, and KVM virtualization technologies.Strong Windows Server administration experience, including exposure to Citrix VDI environments.Experience supporting enterprise server hardware, including HPE ProLiant, Dell PowerEdge, and similar platforms.Proficiency with monitoring and observability tools such as Zabbix, Grafana, or equivalent platforms.Strong scripting and automation skills using Shell scripting and Python.Strong understanding of LAN/WAN networking, including switches, routers, and associated network technologies.Good knowledge of cybersecurity requirements, network security, firewalls, encryption, and HPC security best practices.Familiarity with network monitoring tools and advanced troubleshooting techniques.Strong communication, stakeholder management, and negotiation skills, with the ability to work effectively with customers, vendors, and internal teams.Experience creating and maintaining technical diagrams and documentation using Visio, Draw.io, or similar tools.Familiarity with reservoir engineering simulators and petroleum software, including tNavigator, Nexus-VIP, Schlumberger Eclipse & Intersect, ExxonMobil IMPOWER, Total PUMA, or similar platforms, is an advantage.What working at Core42 offersWith a diverse team of 1,100+ employees from 68 nationalities, we foster an inclusive, innovative and collaborative environment. At Core42, we foster a culture grounded in trust, accountability and high performance. We are united by our values: Grit, where we overcome challenges with resilience and determination, Passion, which drives us to pursue excellence in everything we do, and Impact, as we aim to inspire progress and create meaningful change. Our team members thrive in an environment where each person’s contributions propel us forward, and together, we commit to achieving extraordinary results.Competitive Salary: We offer an attractive salary package based on your skills and experience.Yearly Bonus: In recognition of your contributions, you will receive a performance-based annual bonus.Exclusive Discount Cards: Access special benefits with Esaad and Fazaa cards, offering discounts across a wide range of services.Premium Family Insurance: We provide comprehensive health coverage, including dental, vision, and life insurance, ensuring the well-being of you and your family.Learning & Development: We offer access to top-tier learning platforms to help you grow in your career. Learn at your own pace with unlimited access to premium courses.