IT Administrator — HPC & Data Center Operations
Prepaire Labs · Abu Dhabi, Abu Dhabi Emirate, United Arab Emirates
Apply & track with Apply EdgeJob Description: IT Administrator — HPC & Data Center OperationsDepartment: Information Technology / InfrastructureReports To: Head of IT / Chief Technology Officer Location: Abu Dhabi - On-site (Data Center presence required) Employment Type: Full-TimeAbout the RolePrepaire Labs is seeking an experienced IT Administrator to manage andmaintain our High-Performance Computing (HPC) environment, including CPUand GPU data center infrastructure that powers our AI-driven drug discovery andhealthcare research platforms. The successful candidate will be responsible forend-to-end administration of compute clusters, networking connectivity,healthcare and scientific software licensing, and day-to-day IT operations,ensuring high availability, security, and regulatory compliance across all systemsKey Responsibilities1. HPC & Data Center Administration (CPU/GPU)• Administer, monitor, and maintain HPC clusters comprising CPU and GPUcompute nodes (e.g., NVIDIA H100/A100/L40S, AMD EPYC, Intel Xeonplatforms).• Deploy, configure, and manage cluster workload managers and jobschedulers (SLURM, PBS, or Kubernetes with GPU orchestration).• Manage GPU drivers, CUDA toolkits, container runtimes (Docker,Singularity/Apptainer, NVIDIA Container Toolkit), and ML frameworkssupporting research workloads.• Oversee data center operations: rack and stack, power and coolingmonitoring (PDU/UPS), cable management, hardware lifecycle, and capacityplanning.• Perform preventive maintenance, firmware/BIOS updates, hardwarediagnostics, and coordinate vendor RMA/support cases (NVIDIA, Dell, HPE,Supermicro, Lenovo, etc.).• Manage high-performance storage systems (NFS, Lustre, BeeGFS, Ceph, orenterprise NAS/SAN) and data backup/disaster recovery strategies.• Monitor cluster health, utilization, and performance using tools such asGrafana, Prometheus, Zabbix, Nagios, or NVIDIA DCGM; optimize resourceallocation for research teams.2. Networking & ConnectivityDesign, configure, and maintain LAN/WAN, VLANs, firewalls, VPNs, andwireless infrastructure across office and data center environments.• Administer high-speed interconnects for HPC workloads (InfiniBand, RoCE,10/25/40/100GbE).• Manage core network equipment (Cisco, Juniper, Arista, Fortinet, Mellanox/NVIDIA networking) including switching, routing, and firmware updates.• Ensure secure remote access, site-to-site connectivity, and redundancy/failover for critical links.• Monitor network performance, troubleshoot latency/bandwidth issues, andmaintain network documentation, IP address management (IPAM), andtopology diagrams.• Implement network segmentation and access controls appropriate forsensitive healthcare and research data.3. Healthcare Network, Software & Licensing Management• Manage procurement, deployment, renewal, and compliance of softwarelicenses, including scientific, healthcare, and laboratory applications (e.g.,LIMS, bioinformatics suites, molecular modeling tools, EHR/EMRintegrations where applicable).• Administer license servers (FlexLM/FlexNet, RLM, etc.) for engineering andscientific software.• Track license utilization, forecast needs, and optimize licensing costs acrossteams.• Ensure all systems handling healthcare or patient-related data comply withapplicable regulations and standards (HIPAA, GDPR, ISO 27001, GxP/21 CFRPart 11 as relevant).• Liaise with healthcare technology vendors, managed service providers, andregulatory/compliance teams.4. Systems & Security Administration• Administer Linux (RHEL/Rocky/Ubuntu) and Windows Serverenvironments, including Active Directory / LDAP, DNS, DHCP, and identity/access management (SSO, MFA).• Implement and enforce security policies: patch management, endpointprotection, vulnerability scanning, intrusion detection, and audit logging.• Manage virtualization and cloud resources (VMware, Proxmox, Hyper-V;AWS/Azure/GCP hybrid connectivity where required).• Maintain robust backup, replication, and disaster recovery procedures;conduct periodic DR testing.• Support incident response, root cause analysis, and change management processes.5. IT Operations & User Support• Provide Tier 2/3 support for researchers, scientists, and staff on workstations,peripherals, collaboration tools, and access to HPC resources.• Onboard/offboard users, manage accounts, permissions, and researchproject allocations on compute clusters.• Maintain accurate documentation: SOPs, runbooks, asset inventory, andconfiguration records.• Train end users on responsible and efficient use of HPC and IT resources.• Participate in on-call rotation for critical infrastructure issues.Required Qualifications• Bachelor's degree in Computer Science, Information Technology,Engineering, or equivalent practical experience.• 5+ years of experience in IT systems administration, with 2+ yearsmanaging HPC and/or GPU data center environments.• Strong hands-on expertise in Linux administration (RHEL/CentOS/Rocky/Ubuntu) and shell scripting (Bash, Python).• Proven experience with GPU computing stacks (NVIDIA drivers, CUDA,DCGM) and job schedulers (SLURM preferred).• Solid networking knowledge: TCP/IP, VLANs, routing, firewalls, VPNs, andhigh-speed interconnects (InfiniBand/100GbE a plus).• Experience managing software licensing, license servers, and vendorcontracts.• Familiarity with healthcare data compliance requirements (HIPAA, GDPR)and IT security best practices.• Strong troubleshooting, documentation, and communication skills.Preferred Qualifications• Certifications such as RHCSA/RHCE, CCNA/CCNP, CompTIA Server+/Security+, NVIDIA-certified (e.g., NCA/NCP for AI Infrastructure), VMwareVCP, or ITIL Foundation.• Experience supporting biotech, pharma, healthcare, or research laboratoryenvironments (LIMS, bioinformatics pipelines).• Experience with infrastructure-as-code and automation (Ansible,Terraform, Git).• Exposure to container orchestration (Kubernetes, Helm) and MLOpsenvironments.• Knowledge of data center standards (Tier classifications, hot/cold aisledesign, DCIM tools).• Experience with hybrid cloud HPC bursting and cost optimization.Key Competencies• Ownership mindset with the ability to manage critical 24/7 infrastructure.• Analytical problem-solver who thrives in fast-paced R&D environments.• Detail-oriented approach to compliance, documentation, and changecontrol.• Collaborative communicator able to translate technical matters for scientificand business stakeholders.• Proactive about capacity planning, performance tuning, and continuousimprovement.What We Offer• Opportunity to work at the intersection of AI, HPC, and healthcareinnovation.• Competitive salary and benefits package.• Professional development, training, and certification support.• A collaborative, mission-driven research environment.