Apply Edge Start your job search

Infrastructure Engineer (Dubai)

Connect Tech+Talent · Dubai, United Arab Emirates

Apply & track with Apply Edge

Infrastructure Engineer (SRE + AIOps) Location: Dubai, UAE (Onsite)Duration: 12 Months (Contract)Experience: 4–8 YearsVisa Status: Work Permit, Dependent Visa, or Tourist VisaLocation Requirement: Candidates currently based in Dubai only.Job SummaryWe are seeking a skilled Infrastructure Engineer (SRE + AIOps) with strong experience in cloud infrastructure, site reliability engineering, automation, and AI-driven IT operations. The ideal candidate will have hands-on expertise in Kubernetes, observability, Infrastructure as Code (IaC), and incident management to build and maintain highly available, scalable, and self-healing distributed systems.Key ResponsibilitiesDesign, deploy, and maintain reliable cloud and edge infrastructure for distributed applications.Manage Kubernetes (k3s) environments using Helm and ArgoCD for automated deployments and GitOps workflows.Develop automation scripts and platform tools using Python (FastAPI), Go, or Node.js.Implement monitoring, logging, and observability solutions using Prometheus, Grafana, OpenTelemetry, and ELK/OpenSearch.Apply AIOps techniques for anomaly detection, alert correlation, predictive incident management, and automated remediation.Build and support event-driven systems using Kafka or RabbitMQ.Automate infrastructure provisioning and configuration using Terraform, Helm, and GitOps.Define and monitor SLIs/SLOs to improve system reliability, performance, and availability.Lead incident response, root cause analysis, and postmortems to prevent recurring issues.Develop self-healing capabilities to minimize downtime and manual intervention.Required Skills & Qualifications4–8 years of experience in Infrastructure Engineering, SRE, DevOps, or Platform Engineering.Strong hands-on experience with Kubernetes (k3s), Helm, and ArgoCD.Proficiency in Python, Go, or Node.js for infrastructure automation.Experience with Prometheus, Grafana, OpenTelemetry, and ELK/OpenSearch.Knowledge of AIOps, anomaly detection, alert correlation, and predictive incident management.Experience with Kafka or RabbitMQ and distributed systems.Strong understanding of Terraform, Infrastructure as Code, and GitOps.Practical experience with SLOs/SLIs, incident response, troubleshooting, and postmortems.Familiarity with cloud and edge computing architectures, CI/CD pipelines, and self-healing infrastructure.Preferred SkillsExperience implementing AI/ML-driven operational automation.Exposure to marketplace-style application deployments.Strong problem-solving, troubleshooting, and communication skills.Important: Only candidates currently residing in Dubai, UAE, will be considered. Applicants must have a Work Permit, Dependent Visa, or Tourist Visa.