๐๐ฒ๐ฎ๐ฑ ๐ผ๐ณ ๐๐น๐ผ๐๐ฑ ๐๐ป๐ณ๐ฟ๐ฎ๐๐๐ฟ๐๐ฐ๐๐๐ฟ๐ฒ & ๐ฃ๐ฟ๐ผ๐ฑ๐๐ฐ๐๐ถ๐ผ๐ป (๐๐ป๐๐ถ๐ฑ๐ฒ ๐จ๐๐ ๐๐ฎ๐ป๐ฑ๐ถ๐ฑ๐ฎ๐๐ฒ ๐ข๐ป๐น๐)
Suadeo ยท Dubai, Dubai, United Arab Emirates
Apply & track with Apply Edge๐๐ฒ๐ฎ๐ฑ ๐ผ๐ณ ๐๐น๐ผ๐๐ฑ ๐๐ป๐ณ๐ฟ๐ฎ๐๐๐ฟ๐๐ฐ๐๐๐ฟ๐ฒ & ๐ฃ๐ฟ๐ผ๐ฑ๐๐ฐ๐๐ถ๐ผ๐ป1. Role OverviewThe ๐๐ฒ๐ฎ๐ฑ ๐ผ๐ณ ๐๐น๐ผ๐๐ฑ ๐๐ป๐ณ๐ฟ๐ฎ๐๐๐ฟ๐๐ฐ๐๐๐ฟ๐ฒ & Production is a senior technology leadership role responsible for transforming Suadeo Cloud from a reactive infrastructure/support function into a proactive, world-class reliability engineering and production organization. The role owns the operational health of Suadeoโs customer-facing production environments across multiple countries and is accountable for availability, reliability, performance, security, resilience, scalability and operational quality. The mandate extends beyond traditional Infrastructure/DevOps management to include the production operating model, governance, engineering standards, team capability and continuous improvement. The objective is to build a reliable, industrialized, observable, secure and scalable production organization while reducing operational risk, manual intervention, technical debt and dependency on individual knowledge. 2. ๐ฆ๐๐ฟ๐ฎ๐๐ฒ๐ด๐ถ๐ฐ ๐ฃ๐ฟ๐ถ๐ผ๐ฟ๐ถ๐๐ถ๐ฒ๐๐๐ฒ๐๐ฒ๐น๐ผ๐ฝ ๐ฎ ๐๐ฒ๐ฐ๐๐ฟ๐ฒ, ๐๐ฐ๐ฎ๐น๐ฎ๐ฏ๐น๐ฒ ๐ฎ๐ป๐ฑ ๐ฟ๐ฒ๐๐ถ๐น๐ถ๐ฒ๐ป๐ ๐ฐ๐น๐ผ๐๐ฑ ๐๐๐ฟ๐ฎ๐๐ฒ๐ด๐ aligned with business and product growth.๐ ๐ผ๐ฑ๐ฒ๐ฟ๐ป๐ถ๐๐ฒ ๐บ๐๐น๐๐ถ-๐ฐ๐น๐ผ๐๐ฑ/๐บ๐๐น๐๐ถ-๐ฟ๐ฒ๐ด๐ถ๐ผ๐ป ๐ฒ๐ป๐๐ถ๐ฟ๐ผ๐ป๐บ๐ฒ๐ป๐๐ ๐ฎ๐ป๐ฑ ๐ฟ๐ฒ๐ฑ๐๐ฐ๐ฒ ๐๐ฒ๐ฐ๐ต๐ป๐ถ๐ฐ๐ฎ๐น ๐ฑ๐ฒ๐ฏ๐ and unnecessary complexity.๐๐๐๐ฎ๐ฏ๐น๐ถ๐๐ต ๐๐๐ฟ๐ผ๐ป๐ด ๐ฒ๐ป๐ด๐ถ๐ป๐ฒ๐ฒ๐ฟ๐ถ๐ป๐ด ๐๐๐ฎ๐ป๐ฑ๐ฎ๐ฟ๐ฑ๐, reference architectures and controlled production environments.๐ง๐ฟ๐ฎ๐ป๐๐ถ๐๐ถ๐ผ๐ป ๐ณ๐ฟ๐ผ๐บ ๐ฟ๐ฒ๐ฎ๐ฐ๐๐ถ๐๐ฒ ๐๐๐ฝ๐ฝ๐ผ๐ฟ๐ ๐๐ผ ๐ฝ๐ฟ๐ผ๐ฎ๐ฐ๐๐ถ๐๐ฒ ๐ฆ๐ฅ๐ ๐ฝ๐ฟ๐ฎ๐ฐ๐๐ถ๐ฐ๐ฒ๐ using SLAs, SLOs, SLIs and measurable reliability outcomes.๐๐ฟ๐ถ๐๐ฒ ๐๐ป๐ณ๐ฟ๐ฎ๐๐๐ฟ๐๐ฐ๐๐๐ฟ๐ฒ ๐ฎ๐ ๐๐ผ๐ฑ๐ฒ, ๐๐ถ๐๐ข๐ฝ๐, ๐๐/๐๐ ๐ฎ๐ป๐ฑ ๐ฎ๐๐๐ผ๐บ๐ฎ๐๐ถ๐ผ๐ป across the production lifecycle.๐๐๐๐ฎ๐ฏ๐น๐ถ๐๐ต ๐ฑ๐ถ๐๐ฐ๐ถ๐ฝ๐น๐ถ๐ป๐ฒ๐ฑ ๐ฝ๐ฟ๐ผ๐ฑ๐๐ฐ๐๐ถ๐ผ๐ป ๐ด๐ผ๐๐ฒ๐ฟ๐ป๐ฎ๐ป๐ฐ๐ฒ covering access, change, risk, documentation and operational ownership.๐๐๐ถ๐น๐ฑ ๐ฎ๐ป๐ฑ ๐ฐ๐ผ๐ป๐๐ถ๐ป๐๐ผ๐๐๐น๐ ๐๐ฒ๐๐ ๐๐ถ๐๐ฎ๐๐๐ฒ๐ฟ ๐ฅ๐ฒ๐ฐ๐ผ๐๐ฒ๐ฟ๐ ๐ฎ๐ป๐ฑ ๐๐๐๐ถ๐ป๐ฒ๐๐ ๐๐ผ๐ป๐๐ถ๐ป๐๐ถ๐๐ capabilities against defined RTO/RPO.๐๐บ๐ฏ๐ฒ๐ฑ ๐๐ฒ๐ฐ๐๐ฟ๐ถ๐๐, ๐น๐ฒ๐ฎ๐๐ ๐ฝ๐ฟ๐ถ๐๐ถ๐น๐ฒ๐ด๐ฒ, ๐๐ฒ๐ด๐บ๐ฒ๐ป๐๐ฎ๐๐ถ๐ผ๐ป, ๐๐๐น๐ป๐ฒ๐ฟ๐ฎ๐ฏ๐ถ๐น๐ถ๐๐ ๐บ๐ฎ๐ป๐ฎ๐ด๐ฒ๐บ๐ฒ๐ป๐ ๐ฎ๐ป๐ฑ ๐๐ฒ๐๐ฆ๐ฒ๐ฐ๐ข๐ฝ๐ into production operations.๐๐๐ถ๐น๐ฑ ๐ฎ ๐ต๐ถ๐ด๐ต-๐ฝ๐ฒ๐ฟ๐ณ๐ผ๐ฟ๐บ๐ถ๐ป๐ด ๐๐ป๐ณ๐ฟ๐ฎ๐๐๐ฟ๐๐ฐ๐๐๐ฟ๐ฒ, ๐๐น๐ผ๐๐ฑ, ๐๐ฒ๐๐ข๐ฝ๐ ๐ฎ๐ป๐ฑ ๐ฆ๐ฅ๐ ๐ผ๐ฟ๐ด๐ฎ๐ป๐ถ๐๐ฎ๐๐ถ๐ผ๐ป with clear accountability. 3. ๐๐ฒ๐ ๐ฅ๐ฒ๐๐ฝ๐ผ๐ป๐๐ถ๐ฏ๐ถ๐น๐ถ๐๐ถ๐ฒ๐๐๐น๐ผ๐๐ฑ & ๐๐ป๐ณ๐ฟ๐ฎ๐๐๐ฟ๐๐ฐ๐๐๐ฟ๐ฒ ๐ฆ๐๐ฟ๐ฎ๐๐ฒ๐ด๐๐ข๐๐ป ๐๐ต๐ฒ ๐ฐ๐น๐ผ๐๐ฑ ๐ถ๐ป๐ณ๐ฟ๐ฎ๐๐๐ฟ๐๐ฐ๐๐๐ฟ๐ฒ ๐ฟ๐ผ๐ฎ๐ฑ๐บ๐ฎ๐ฝ ๐ฎ๐ฐ๐ฟ๐ผ๐๐ ๐๐ช๐ฆ, ๐๐๐๐ฟ๐ฒ ๐ฎ๐ป๐ฑ/๐ผ๐ฟ ๐๐๐ฃ.Define scalable, secure, highly available and cost-effective cloud architecture and standards.Lead infrastructure modernization, technical-debt reduction, automation and multi-region/HA initiatives.Establish Infrastructure as Code and configuration-management standards. ๐ฃ๐ฟ๐ผ๐ฑ๐๐ฐ๐๐ถ๐ผ๐ป ๐ข๐๐ป๐ฒ๐ฟ๐๐ต๐ถ๐ฝ & ๐ฅ๐ฒ๐น๐ถ๐ฎ๐ฏ๐ถ๐น๐ถ๐๐ยท Own end-to-end production health across customer environments, including Kubernetes, databases, storage, certificates, backups, capacity and critical dependencies.Establish proactive SRE practices, production readiness standards, operational ownership and runbooks.Ensure production environments are observable, supportable, secure and continuously improved.Drive availability, reliability and performance improvements for business-critical services.๐๐ป๐ฐ๐ถ๐ฑ๐ฒ๐ป๐, ๐ฃ๐ฟ๐ผ๐ฏ๐น๐ฒ๐บ & ๐๐ต๐ฎ๐ป๐ด๐ฒ ๐ ๐ฎ๐ป๐ฎ๐ด๐ฒ๐บ๐ฒ๐ป๐ยท Establish ITIL-aligned incident, problem, change and escalation processes.ยท Lead major incidents and ensure effective on-call, severity and escalation models.ยท Own MTTA, MTTR, incident recurrence and post-incident improvement.ยท Ensure root-cause analysis is completed, and corrective actions are tracked to closure.ยท Strengthen change and release governance to reduce production risk.๐ข๐ฏ๐๐ฒ๐ฟ๐๐ฎ๐ฏ๐ถ๐น๐ถ๐๐ & ๐ข๐ฝ๐ฒ๐ฟ๐ฎ๐๐ถ๐ผ๐ป๐ฎ๐น ๐๐ป๐๐ฒ๐น๐น๐ถ๐ด๐ฒ๐ป๐ฐ๐ฒยท Own monitoring, logging, metrics, tracing, alerting and operational dashboards.ยท Establish meaningful SLAs, SLOs and SLIs and use them to drive reliability decisions.ยท Reduce alert fatigue and improve proactive detection of performance, availability and capacity issues. ๐๐ฒ๐๐ข๐ฝ๐, ๐๐/๐๐ & ๐๐๐๐ผ๐บ๐ฎ๐๐ถ๐ผ๐ปยท Drive CI/CD, Terraform/IaC, GitOps and automated deployment practices.ยท Promote reliable deployment strategies including zero-downtime, blue/green and canary releases where appropriate.ยท Automate repetitive operational tasks and eliminate avoidable manual intervention.ยท Partner with R&D to improve release readiness, deployment quality and production stability. ๐ฅ๐ฒ๐๐ถ๐น๐ถ๐ฒ๐ป๐ฐ๐ฒ, ๐๐ฅ & ๐๐๐๐ถ๐ป๐ฒ๐๐ ๐๐ผ๐ป๐๐ถ๐ป๐๐ถ๐๐ยท Design and maintain HA, backup, replication and failover strategies.ยท Define and validate RTO/RPO requirements.ยท Conduct regular backup/restore tests, failover exercises and DR simulations.ยท Ensure recoverability is demonstrated, documented and continuously improved. ๐ฆ๐ฒ๐ฐ๐๐ฟ๐ถ๐๐, ๐ฅ๐ถ๐๐ธ & ๐๐ผ๐บ๐ฝ๐น๐ถ๐ฎ๐ป๐ฐ๐ฒยท Embed secure-by-design practices across cloud and production environments.ยท Govern IAM, least privilege, network segmentation, secrets, certificates, patching and vulnerability remediation.ยท Maintain security logging, hardening and operational controls.ยท Support audits, penetration testing, customer security assessments and regulatory/compliance requirements. ๐๐ฎ๐ฝ๐ฎ๐ฐ๐ถ๐๐, ๐ฃ๐ฒ๐ฟ๐ณ๐ผ๐ฟ๐บ๐ฎ๐ป๐ฐ๐ฒ & ๐๐น๐ผ๐๐ฑ ๐๐ผ๐๐ยท Establish proactive capacity planning, forecasting and performance management.ยท Identify infrastructure and application bottlenecks and coordinate improvements with R&D.ยท Drive cloud cost optimization and FinOps discipline without compromising reliability or security. ๐๐ฒ๐ฎ๐ฑ๐ฒ๐ฟ๐๐ต๐ถ๐ฝ, ๐ข๐ฟ๐ด๐ฎ๐ป๐ถ๐๐ฎ๐๐ถ๐ผ๐ป & ๐ง๐ฎ๐น๐ฒ๐ป๐ยท Assess and evolve the Infrastructure/Cloud/DevOps/SRE operating model, roles and accountability.ยท Build a capable, scalable team with clear ownership, on-call and escalation structures.ยท Recruit, mentor and develop technical talent and strengthen operational discipline.ยท Provide hands-on technical leadership when complex production issues require intervention.๐ฅ&๐ & ๐๐ฟ๐ผ๐๐-๐๐๐ป๐ฐ๐๐ถ๐ผ๐ป๐ฎ๐น ๐ฃ๐ฎ๐ฟ๐๐ป๐ฒ๐ฟ๐๐ต๐ถ๐ฝEstablish production-readiness requirements for new products, releases and customer environments.Lead cross-functional investigations across infrastructure, networking, Kubernetes, databases, storage, security and application architecture.Ensure production engineering is treated as a reliability and engineering functionโnot simply a ticket-routing/support function. ๐๐ป๐ผ๐๐น๐ฒ๐ฑ๐ด๐ฒ, ๐๐ผ๐ฐ๐๐บ๐ฒ๐ป๐๐ฎ๐๐ถ๐ผ๐ป & ๐๐ผ๐๐ฒ๐ฟ๐ป๐ฎ๐ป๐ฐ๐ฒMaintain accurate architecture diagrams, infrastructure inventories, dependency maps and operational documentation.Establish and maintain runbooks for incidents, deployments, escalation, backup/restore and DR.Reduce dependency on individual knowledge through structured documentation and knowledge transfer.Ensure operational decisions and controls are traceable and audit-ready. ๐๐ ๐ฒ๐ฐ๐๐๐ถ๐๐ฒ ๐ฅ๐ฒ๐ฝ๐ผ๐ฟ๐๐ถ๐ป๐ด & ๐๐ฃ๐๐ยท Provide executive-level visibility on:ยท Availability, SLA/SLO performance and reliability trends.ยท Critical incidents, MTTA, MTTR and recurring issues.ยท Deployment/change success and production stability.ยท Capacity, cloud cost and operational efficiency.ยท Backup/restore and DR readiness.ยท Vulnerabilities, critical risks and technical debt.ยท Post-mortem actions and overall production health. 4. ๐ง๐ฒ๐ฐ๐ต๐ป๐ถ๐ฐ๐ฎ๐น ๐๐ป๐๐ถ๐ฟ๐ผ๐ป๐บ๐ฒ๐ป๐ & ๐๐ ๐ฝ๐ฒ๐ฐ๐๐ฒ๐ฑ ๐๐ฒ๐ฝ๐๐ตยท Strong practical knowledge across:ยท AWS, Azure and/or GCP; multi-cloud/multi-region environments.ยท Kubernetes/OpenShift, Docker, Helm and ArgoCD/GitOps.ยท CI/CD, Terraform and Infrastructure as Code.ยท Linux, networking, DNS, TLS, reverse proxies and load balancing.ยท SQL databases, S3/object storage, Redis and Kafka.ยท Prometheus, Grafana, centralized logging, APM and distributed tracing.ยท Backup, restore, DR, HA and resilience engineering.ยท IAM, OpenID, LDAP/AD and cloud security controls.ยท APIs, distributed systems and application performance. 5. ๐๐ฎ๐ป๐ฑ๐ถ๐ฑ๐ฎ๐๐ฒ ๐ฃ๐ฟ๐ผ๐ณ๐ถ๐น๐ฒ & ๐ค๐๐ฎ๐น๐ถ๐ณ๐ถ๐ฐ๐ฎ๐๐ถ๐ผ๐ป๐12+ years of experience across IT Infrastructure, Cloud, DevOps, SRE or Production Operations.5+ years in senior technology leadership with ownership of critical, customer-facing production environments.Proven experience leading teams through major incidents, operational transformation and production governance.Strong understanding of SRE, cloud operating models, CI/CD, automation and ITIL-aligned practices.Demonstrated ability to balance executive-level strategy with hands-on technical intervention when required.Strategic, structured, proactive and highly accountable, with strong risk and operational discipline.Bachelorโs/Masterโs degree in Computer Science, Engineering, IT or a related field; MBA/Technology Management is an advantage. 6. ๐ฃ๐ฟ๐ฒ๐ณ๐ฒ๐ฟ๐ฟ๐ฒ๐ฑ ๐๐ฒ๐ฟ๐๐ถ๐ณ๐ถ๐ฐ๐ฎ๐๐ถ๐ผ๐ป๐AWS Solutions Architect Professional, Google Professional Cloud Architect, Azure Solutions Architect Expert, CKA, Terraform Associate, advanced SRE/DevOps certifications, CISSP, CISM and/or ITIL Managing Professional.