Apply Edge Start your job search

Lead Engineer – Infrastructure & Cloud Engineering

SmartChoice International GCC · Abu Dhabi, Abu Dhabi Emirate, United Arab Emirates

Apply & track with Apply Edge
Lead Engineer – Infrastructure & Cloud EngineeringAbu Dhabi, UAE | PermanentAbout the ClientOur client is a high-growth technology powerhouse operating at the forefront of AI, cloud and digital infrastructure. With an ambitious vision for the future of technology, the organization is bringing together world-class engineering talent, advanced infrastructure and next-generation platforms to solve some of the most complex technology challenges at scale.This is an opportunity to join an environment where deep engineering, ambitious technology and genuine innovationcome together, with the chance to influence the architecture of critical platforms and help shape what comes next.The OpportunityWe are looking for an experienced Lead Engineer – Infrastructure & Cloud Engineering to provide technical leadership across private cloud, virtualization, container and observability platforms.The role combines deep hands-on engineering expertise with technical leadership, helping shape platform strategy, engineering standards, operational excellence and the adoption of emerging technologies.You will work closely with architecture, security, platform engineering, SRE and operations teams to build, operate and continuously improve highly available infrastructure platforms supporting demanding technology environments.This is a hands-on technical leadership role for someone who enjoys solving complex infrastructure challenges, leading engineering initiatives and helping teams adopt modern approaches to cloud, automation and platform engineering.Key ResponsibilitiesLead the technical design, implementation and lifecycle management of large-scale private cloud, virtualization and container platforms, including OpenStack and OpenShift or comparable technologies.Define and drive engineering standards, design principles, operational best practices and technical governance across infrastructure and cloud platforms.Provide technical leadership and mentorship to engineering teams, supporting technical decision-making, complex problem-solving, knowledge sharing and continuous development.Lead the strategy, design and adoption of observability platforms, covering metrics, logs, traces, dashboards, alerting, service health and SLO/SLI capabilities.Drive the adoption of AI-assisted operations and intelligent automation to improve operational efficiency, incident management, root cause analysis, platform reliability and service automation.Lead the evaluation, integration and adoption of new technologies, products and architectural approaches in collaboration with Architecture, Product Engineering, SRE, Security and Operations teams.Provide technical oversight for complex platform upgrades, migrations, production changes and infrastructure transformation initiatives.Act as a senior technical escalation point for critical production incidents, leading complex troubleshooting, incident response, root cause analysis and service recovery.Drive capacity planning, scalability, performance optimisation, resilience engineering and operational readiness across infrastructure platforms.Ensure observability, automation and operational tooling are effectively integrated with service management and incident management processes.Work closely with security, compliance and risk teams to ensure infrastructure platforms meet appropriate cybersecurity, governance and regulatory requirements.Define and champion automation strategies using Infrastructure as Code, GitOps, CI/CD and platform engineering practices.Lead the development and maintenance of technical standards, architecture documentation, operational procedures, design guides and knowledge resources.Collaborate with leadership teams on infrastructure strategy, roadmap planning, technology evaluation and long-term platform evolution.What We're Looking ForBachelor's or Master's degree in Computer Science, Engineering, Software Engineering or a related technology discipline, or equivalent practical experience.8+ years of experience designing, implementing, operating, troubleshooting and leading large-scale cloud, infrastructure or platform engineering environments.Strong hands-on experience with private cloud and virtualization platforms, with proven experience in OpenStack, OpenShift or comparable technologies.Experience working with container platforms and orchestration technologies, with a good understanding of containerised infrastructure.Proven experience providing technical leadership across complex infrastructure engineering environments.Strong understanding of compute infrastructure, including x86/ARM architecture, Linux, KVM, virtualization, server hardware, firmware and infrastructure lifecycle management.Expert-level Linux administration and troubleshooting skills, including performance analysis, system tuning, patching and operational support.Strong understanding of data centre networking, including TCP/IP, routing, firewalls, load balancing, VLAN/VXLAN, DNS, DHCP and related technologies.Deep understanding of observability principles, including metrics, logs, traces, alerting, dashboards and service-level monitoring.Hands-on experience with modern observability technologies such as Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch or comparable platforms.Proven experience operating large-scale private or public cloud environments, highly available infrastructure or mission-critical technology platforms.Strong experience with Infrastructure as Code, automation, CI/CD and GitOps, using technologies such as Terraform, Ansible, Helm or comparable tools.Strong scripting or programming capabilities using Python, Go, Bash or similar languages.Experience with AI-assisted operations, workflow automation or intelligent infrastructure management is highly desirable.Strong understanding of software-defined infrastructure, platform engineering, reliability engineering, incident management and infrastructure lifecycle automation.Experience integrating observability and operational tooling with service management and incident management processes.Understanding of infrastructure security, security monitoring, compliance, data governance and risk management.Relevant certifications across cloud, Linux, virtualization, Kubernetes, OpenStack, OpenShift or IT service management are advantageous.Excellent analytical, troubleshooting, communication, documentation, stakeholder management, mentoring and technical leadership skills.You'll be a technically strong infrastructure engineer who is comfortable operating at both hands-on engineering and technical leadership level.You'll be able to get into the detail of complex infrastructure problems while also stepping back to define standards, influence technical decisions and guide engineering teams.You should have a strong interest in cloud infrastructure, automation, observability and emerging technology, with the ability to identify where new approaches — including AI-assisted operations — can genuinely improve reliability and operational efficiency.