Apply Edge Start your job search

Lead Engineer - Infrastructure & Cloud Engineering

Core42 · Abu Dhabi, Abu Dhabi Emirate, United Arab Emirates

Apply & track with Apply Edge
Lead Engineer - Infrastructure & Cloud EngineeringAbout UsCore42, a leader in AI-powered cloud and digital infrastructure, is driving transformative technology solutions globally. Leveraging advanced resources and partnerships, Core42 empowers clients to harness sovereign AI infrastructure, especially in sectors with stringent regulatory needs. With a mission to redefine digital transformation, we combine sovereign capabilities with scalable, high-performance compute infrastructure, positioning itself at the forefront of AI innovation in the Middle East and beyond.Role The Lead Engineer – Infrastructure & Cloud Engineering provides technical leadership for the design, implementation, operation, and continuous evolution of large-scale private cloud, virtualization, container, and observability platforms across Core42 infrastructure environments.The role combines deep hands-on technical expertise with technical leadership responsibilities, driving platform strategy, architecture alignment, engineering standards, operational excellence, and adoption of emerging technologies across Infrastructure Engineering teams.Responsibilities Lead the technical design, architecture, implementation, and lifecycle management of large-scale private cloud, virtualization, container, and observability platforms based on OpenStack, Red Hat OpenShift, and related infrastructure technologies.Define and drive engineering standards, design principles, operational best practices, and technical governance frameworks across infrastructure and cloud platforms.Provide technical leadership and mentorship to engineering teams, supporting technical decision-making, problem-solving, knowledge sharing, and continuous skill development.Lead the strategy, design, and adoption of observability platforms and services, including metrics, logs, traces, dashboards, alerting, SLOs/SLIs, and service health visibility capabilities.Drive the adoption of AI-assisted operations and Agentic AI technologies to improve operational efficiency, incident management, root cause analysis, platform reliability, and service automation.Lead the evaluation, integration, and adoption of new technologies, products, and architectural approaches in collaboration with Architecture, Product Engineering, SRE, Security, and Operations teams.Own technical oversight of complex platform upgrades, migrations, production changes, incident response activities, root cause analysis, and reliability improvement initiatives.Drive capacity planning, scalability strategies, performance optimization, resiliency engineering, and operational readiness activities across cloud and infrastructure platforms.Ensure observability, automation, operational tooling, and platform services are effectively integrated with ITSM, service management, incident management, and operational governance processes.Partner with security, compliance, and risk management teams to ensure infrastructure platforms meet cybersecurity, governance, regulatory, and data sovereignty requirements.Define and champion automation strategies using Infrastructure as Code, GitOps, CI/CD, platform engineering, and infrastructure lifecycle management best practices.Lead the development and maintenance of technical standards, architecture documentation, operational procedures, design guides, and knowledge management repositories.Serve as a senior technical escalation point for critical production incidents and participate in on-call duty rotations, leading complex troubleshooting, incident response coordination, root cause analysis, and service recovery activities.Collaborate with leadership teams on infrastructure strategy, roadmap planning, investment decisions, and long-term platform evolution.Qualification, Experience, Competence and Certifications Bachelor’s or Master’s degree in Computer Science, Engineering, Software Engineering, or a related technology discipline; or equivalent practical experience.8+ years of experience designing, implementing, operating, troubleshooting, and leading large-scale private cloud, virtualization, infrastructure, observability, or platform engineering environments.Deep hands-on expertise and technical leadership experience with at least one major cloud or virtualization platform, such as OpenStack, Proxmox, Red Hat OpenShift, or equivalent technologies.Experience driving the adoption of AI-assisted operations, workflow automation, and integration of observability and ITSM platforms to improve operational efficiency, incident response, and service reliability.Strong expertise in compute technologies, including x86 hardware, KVM, Linux operating systems, hypervisors, firmware, server lifecycle management, and orchestration services.Strong hands-on experience with at least one major cloud or virtualization platform, such as OpenStack, Proxmox, Red Hat OpenShift, or equivalent technologies.Hands-on experience with AI-assisted operations, workflow automation, and integration of observability and ITSM platforms to improve operational efficiency, incident response, and service reliability.Strong hands-on experience with compute technologies, including x86 hardware, KVM, Linux operating systems, hypervisors, firmware, server lifecycle management, and orchestration services.Expert-level Linux administration and troubleshooting skills, including performance analysis, system tuning, patching, and operational leadership of Linux-based infrastructure environments.Strong understanding of hardware architecture and components, including x86/ARM, NUMA, memory channels, NICs, GPU/AI accelerators, firmware, and large-scale server platforms, with the ability to guide platform design and technology selection.Strong understanding of data center networking concepts, including OSI model, TCP/IP, routing, firewalls, load balancing, VLAN/VXLAN, DNS, DHCP, and related technologies.Deep expertise in observability concepts and platforms, including metrics, logs, traces, dashboards, alerting, SLOs/SLIs, OpenTelemetry, Prometheus, Grafana, ELK/OpenSearch, Splunk, Zabbix, or similar technologies.Proven experience leading large-scale public or private cloud environments, cloud service provider platforms, mission-critical infrastructure, or high-availability managed service environments.Strong experience driving automation strategies, Infrastructure as Code, CI/CD, GitOps, and platform engineering practices using Ansible, Terraform, Helm, Jenkins, GitLab CI/CD, Python, Go, Bash, or similar technologies.Strong understanding of software-defined infrastructure, CI/CD principles, infrastructure lifecycle automation, reliability engineering, incident management, operational governance, and engineering best practices.Understanding of security monitoring, SIEM concepts, compliance requirements, data governance, and risk management considerations in cloud and infrastructure environments.Relevant certifications in Linux, virtualization, cloud computing, OpenStack, Kubernetes, OpenShift, ITSM, or related technologies are advantageous.Excellent analytical, troubleshooting, communication, documentation, stakeholder management, mentoring, and technical leadership skills.What working at Core42 offersWith a diverse team of 1,100+ employees from 68 nationalities, we foster an inclusive, innovative and collaborative environment. At Core42, we foster a culture grounded in trust, accountability and high performance. We are united by our values: Grit, where we overcome challenges with resilience and determination, Passion, which drives us to pursue excellence in everything we do, and Impact, as we aim to inspire progress and create meaningful change. Our team members thrive in an environment where each person’s contributions propel us forward, and together, we commit to achieving extraordinary results.Competitive Salary: We offer an attractive salary package based on your skills and experienceYearly Bonus: In recognition of your contributions, you will receive a performance-based annual bonusExclusive Discount Cards: Access special benefits with Esaad and Fazaa cards, offering discounts across a wide range of servicesPremium Family Insurance: We provide comprehensive health coverage, including dental, vision and life insurance, ensuring the well-being of you and your familyLearning & Development: We offer access to top-tier learning platforms to help you grow in your career. Learn at your own pace with unlimited