Senior Engineer - Infrastructure and Cloud Engineering
Core42 · Abu Dhabi, Abu Dhabi Emirate, United Arab Emirates
Apply & track with Apply EdgeSenior Engineer - Infrastructure and Cloud EngineeringAbout UsCore42, a leader in AI-powered cloud and digital infrastructure, is driving transformative technology solutions globally. Leveraging advanced resources and partnerships, Core42 empowers clients to harness sovereign AI infrastructure, especially in sectors with stringent regulatory needs. With a mission to redefine digital transformation, we combine sovereign capabilities with scalable, high-performance compute infrastructure, positioning itself at the forefront of AI innovation in the Middle East and beyondRole The Senior Engineer – Infrastructure & Cloud Engineering contributes to the design, implementation, operation, and continuous improvement of large-scale private cloud, virtualization, and observability platforms across Core42 infrastructure environments.The role is hands-on and focused on deploying, operating, troubleshooting, automating, and optimizing virtualization and observability platforms capabilities across production and non-production environments.Responsibilities Design, implement, and operate observability platforms and services, including metrics, logs, traces, dashboards, alerting, and service health visibility using technologies such as Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, and related platforms.Contribute to the design, implementation, and operation of large-scale private cloud, virtualization, and container platforms based on OpenStack, Red Hat OpenShift, and related infrastructure technologies.Develop, integrate, and maintain AI-powered operational agents that leverage observability, monitoring, logging, ITSM, and platform telemetry systems to automate incident detection, root cause analysis, operational workflows, and approved remediation activities in accordance with established operational processes and governance controls.Collaborate with architecture, product, platform engineering, SRE, security, and operations teams on technology evaluation, integration, solution design, and operational readiness.Support capacity planning, performance optimization, production upgrades, migrations, incident response, root cause analysis, and continuous improvement of platform reliability and operational resilience.Integrate observability and platform management capabilities with ITSM, incident, change, and problem management processes and tools such as Jira, ServiceNow, PagerDuty, Opsgenie, or similar platforms.Collaborate with security teams to ensure compute, virtualization, cloud, observability, and management platforms are secure, hardened, and aligned with cybersecurity, compliance, and data governance requirements.Create and maintain technical documentation, including implementation guides, operational procedures, dashboards, runbooks, diagrams, standards, and knowledge base articles.Prepare and deliver technical knowledge transfer sessions for operational teams, SREs, and engineering stakeholders.Participate in on-call rotations and provide technical escalation support for critical production incidents, major service disruptions, and platform emergencies, ensuring timely restoration of services and effective root cause resolution.Work with process and operations teams to improve support workflows, service onboarding, operational procedures, and collaboration efficiency.Qualification, Experience, Competence and Certifications Bachelor’s or Master’s degree in Computer Science, Engineering, Software Engineering, or a related technology discipline; or equivalent practical experience.5+ years of hands-on experience designing, implementing, operating, troubleshooting, and managing private cloud, virtualization, infrastructure, observability, or platform engineering environments.Strong hands-on experience with at least one major cloud or virtualization platform, such as OpenStack, Proxmox, Red Hat OpenShift, or equivalent technologies.Hands-on experience with AI-assisted operations, workflow automation, and integration of observability and ITSM platforms to improve operational efficiency, incident response, and service reliability.Strong hands-on experience with compute technologies, including x86 hardware, KVM, Linux operating systems, hypervisors, firmware, server lifecycle management, and orchestration services.Expert-level Linux administration skills, including troubleshooting, performance analysis, system tuning, patching, and operational support of Linux-based infrastructure environments.Strong understanding of hardware architecture and components, including x86/ARM, NUMA, memory channels, NICs, GPU/AI accelerators, firmware, and large-scale server platforms.Good understanding of data center networking concepts, including OSI model, TCP/IP, routing, firewalls, load balancing, VLAN/VXLAN, DNS, DHCP, and related technologies.Good understanding of storage types, architectures, and protocols, including object, block, file storage, and storage integration with cloud platforms.Hands-on experience with observability concepts and platforms, including metrics, logs, traces, dashboards, alerting, SLOs/SLIs, OpenTelemetry, Prometheus, Grafana, ELK/OpenSearch, Splunk, Zabbix, or similar technologies.Experience managing large-scale public or private cloud environments, cloud service provider platforms, mission-critical infrastructure, or high-availability managed service environments is highly desirable.Strong experience with automation, Infrastructure as Code, CI/CD, GitOps, and scripting using Ansible, Terraform, Helm, Jenkins, GitLab CI/CD, Python, Go, Bash, or similar technologies.Understanding of software-defined infrastructure, CI/CD principles, infrastructure lifecycle automation, reliability engineering, incident management, and operational best practices.Understanding of security monitoring, SIEM concepts, compliance requirements, and data governance considerations in cloud and infrastructure environments is an advantage.Relevant certifications in Linux, virtualization, cloud computing, OpenStack, Kubernetes, OpenShift, or ITSM are advantageous.Strong analytical, troubleshooting, communication, documentation, stakeholder management, and problem-solving skills.What working at Core42 offersWith a diverse team of 1,100+ employees from 68 nationalities, we foster an inclusive, innovative and collaborative environment. At Core42, we foster a culture grounded in trust, accountability and high performance. We are united by our values: Grit, where we overcome challenges with resilience and determination, Passion, which drives us to pursue excellence in everything we do, and Impact, as we aim to inspire progress and create meaningful change. Our team members thrive in an environment where each person’s contributions propel us forward, and together, we commit to achieving extraordinary results.Competitive Salary: We offer an attractive salary package based on your skills and experienceYearly Bonus: In recognition of your contributions, you will receive a performance-based annual bonusExclusive Discount Cards: Access special benefits with Esaad and Fazaa cards, offering discounts across a wide range of servicesPremium Family Insurance: We provide comprehensive health coverage, including dental, vision and life insurance, ensuring the well-being of you and your family