Apply Edge Start your job search

๐—›๐—ฒ๐—ฎ๐—ฑ ๐—ผ๐—ณ ๐—–๐—น๐—ผ๐˜‚๐—ฑ ๐—œ๐—ป๐—ณ๐—ฟ๐—ฎ๐˜€๐˜๐—ฟ๐˜‚๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ & ๐—ฃ๐—ฟ๐—ผ๐—ฑ๐˜‚๐—ฐ๐˜๐—ถ๐—ผ๐—ป (๐—œ๐—ป๐˜€๐—ถ๐—ฑ๐—ฒ ๐—จ๐—”๐—˜ ๐—–๐—ฎ๐—ป๐—ฑ๐—ถ๐—ฑ๐—ฎ๐˜๐—ฒ ๐—ข๐—ป๐—น๐˜†)

Suadeo ยท Dubai, Dubai, United Arab Emirates

Apply & track with Apply Edge

๐—›๐—ฒ๐—ฎ๐—ฑ ๐—ผ๐—ณ ๐—–๐—น๐—ผ๐˜‚๐—ฑ ๐—œ๐—ป๐—ณ๐—ฟ๐—ฎ๐˜€๐˜๐—ฟ๐˜‚๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ & ๐—ฃ๐—ฟ๐—ผ๐—ฑ๐˜‚๐—ฐ๐˜๐—ถ๐—ผ๐—ป1. Role OverviewThe ๐—›๐—ฒ๐—ฎ๐—ฑ ๐—ผ๐—ณ ๐—–๐—น๐—ผ๐˜‚๐—ฑ ๐—œ๐—ป๐—ณ๐—ฟ๐—ฎ๐˜€๐˜๐—ฟ๐˜‚๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ & Production is a senior technology leadership role responsible for transforming Suadeo Cloud from a reactive infrastructure/support function into a proactive, world-class reliability engineering and production organization. The role owns the operational health of Suadeoโ€™s customer-facing production environments across multiple countries and is accountable for availability, reliability, performance, security, resilience, scalability and operational quality. The mandate extends beyond traditional Infrastructure/DevOps management to include the production operating model, governance, engineering standards, team capability and continuous improvement. The objective is to build a reliable, industrialized, observable, secure and scalable production organization while reducing operational risk, manual intervention, technical debt and dependency on individual knowledge. 2. ๐—ฆ๐˜๐—ฟ๐—ฎ๐˜๐—ฒ๐—ด๐—ถ๐—ฐ ๐—ฃ๐—ฟ๐—ถ๐—ผ๐—ฟ๐—ถ๐˜๐—ถ๐—ฒ๐˜€๐——๐—ฒ๐˜ƒ๐—ฒ๐—น๐—ผ๐—ฝ ๐—ฎ ๐˜€๐—ฒ๐—ฐ๐˜‚๐—ฟ๐—ฒ, ๐˜€๐—ฐ๐—ฎ๐—น๐—ฎ๐—ฏ๐—น๐—ฒ ๐—ฎ๐—ป๐—ฑ ๐—ฟ๐—ฒ๐˜€๐—ถ๐—น๐—ถ๐—ฒ๐—ป๐˜ ๐—ฐ๐—น๐—ผ๐˜‚๐—ฑ ๐˜€๐˜๐—ฟ๐—ฎ๐˜๐—ฒ๐—ด๐˜† aligned with business and product growth.๐— ๐—ผ๐—ฑ๐—ฒ๐—ฟ๐—ป๐—ถ๐˜‡๐—ฒ ๐—บ๐˜‚๐—น๐˜๐—ถ-๐—ฐ๐—น๐—ผ๐˜‚๐—ฑ/๐—บ๐˜‚๐—น๐˜๐—ถ-๐—ฟ๐—ฒ๐—ด๐—ถ๐—ผ๐—ป ๐—ฒ๐—ป๐˜ƒ๐—ถ๐—ฟ๐—ผ๐—ป๐—บ๐—ฒ๐—ป๐˜๐˜€ ๐—ฎ๐—ป๐—ฑ ๐—ฟ๐—ฒ๐—ฑ๐˜‚๐—ฐ๐—ฒ ๐˜๐—ฒ๐—ฐ๐—ต๐—ป๐—ถ๐—ฐ๐—ฎ๐—น ๐—ฑ๐—ฒ๐—ฏ๐˜ and unnecessary complexity.๐—˜๐˜€๐˜๐—ฎ๐—ฏ๐—น๐—ถ๐˜€๐—ต ๐˜€๐˜๐—ฟ๐—ผ๐—ป๐—ด ๐—ฒ๐—ป๐—ด๐—ถ๐—ป๐—ฒ๐—ฒ๐—ฟ๐—ถ๐—ป๐—ด ๐˜€๐˜๐—ฎ๐—ป๐—ฑ๐—ฎ๐—ฟ๐—ฑ๐˜€, reference architectures and controlled production environments.๐—ง๐—ฟ๐—ฎ๐—ป๐˜€๐—ถ๐˜๐—ถ๐—ผ๐—ป ๐—ณ๐—ฟ๐—ผ๐—บ ๐—ฟ๐—ฒ๐—ฎ๐—ฐ๐˜๐—ถ๐˜ƒ๐—ฒ ๐˜€๐˜‚๐—ฝ๐—ฝ๐—ผ๐—ฟ๐˜ ๐˜๐—ผ ๐—ฝ๐—ฟ๐—ผ๐—ฎ๐—ฐ๐˜๐—ถ๐˜ƒ๐—ฒ ๐—ฆ๐—ฅ๐—˜ ๐—ฝ๐—ฟ๐—ฎ๐—ฐ๐˜๐—ถ๐—ฐ๐—ฒ๐˜€ using SLAs, SLOs, SLIs and measurable reliability outcomes.๐——๐—ฟ๐—ถ๐˜ƒ๐—ฒ ๐—œ๐—ป๐—ณ๐—ฟ๐—ฎ๐˜€๐˜๐—ฟ๐˜‚๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ ๐—ฎ๐˜€ ๐—–๐—ผ๐—ฑ๐—ฒ, ๐—š๐—ถ๐˜๐—ข๐—ฝ๐˜€, ๐—–๐—œ/๐—–๐—— ๐—ฎ๐—ป๐—ฑ ๐—ฎ๐˜‚๐˜๐—ผ๐—บ๐—ฎ๐˜๐—ถ๐—ผ๐—ป across the production lifecycle.๐—˜๐˜€๐˜๐—ฎ๐—ฏ๐—น๐—ถ๐˜€๐—ต ๐—ฑ๐—ถ๐˜€๐—ฐ๐—ถ๐—ฝ๐—น๐—ถ๐—ป๐—ฒ๐—ฑ ๐—ฝ๐—ฟ๐—ผ๐—ฑ๐˜‚๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐—ด๐—ผ๐˜ƒ๐—ฒ๐—ฟ๐—ป๐—ฎ๐—ป๐—ฐ๐—ฒ covering access, change, risk, documentation and operational ownership.๐—•๐˜‚๐—ถ๐—น๐—ฑ ๐—ฎ๐—ป๐—ฑ ๐—ฐ๐—ผ๐—ป๐˜๐—ถ๐—ป๐˜‚๐—ผ๐˜‚๐˜€๐—น๐˜† ๐˜๐—ฒ๐˜€๐˜ ๐——๐—ถ๐˜€๐—ฎ๐˜€๐˜๐—ฒ๐—ฟ ๐—ฅ๐—ฒ๐—ฐ๐—ผ๐˜ƒ๐—ฒ๐—ฟ๐˜† ๐—ฎ๐—ป๐—ฑ ๐—•๐˜‚๐˜€๐—ถ๐—ป๐—ฒ๐˜€๐˜€ ๐—–๐—ผ๐—ป๐˜๐—ถ๐—ป๐˜‚๐—ถ๐˜๐˜† capabilities against defined RTO/RPO.๐—˜๐—บ๐—ฏ๐—ฒ๐—ฑ ๐˜€๐—ฒ๐—ฐ๐˜‚๐—ฟ๐—ถ๐˜๐˜†, ๐—น๐—ฒ๐—ฎ๐˜€๐˜ ๐—ฝ๐—ฟ๐—ถ๐˜ƒ๐—ถ๐—น๐—ฒ๐—ด๐—ฒ, ๐˜€๐—ฒ๐—ด๐—บ๐—ฒ๐—ป๐˜๐—ฎ๐˜๐—ถ๐—ผ๐—ป, ๐˜ƒ๐˜‚๐—น๐—ป๐—ฒ๐—ฟ๐—ฎ๐—ฏ๐—ถ๐—น๐—ถ๐˜๐˜† ๐—บ๐—ฎ๐—ป๐—ฎ๐—ด๐—ฒ๐—บ๐—ฒ๐—ป๐˜ ๐—ฎ๐—ป๐—ฑ ๐——๐—ฒ๐˜ƒ๐—ฆ๐—ฒ๐—ฐ๐—ข๐—ฝ๐˜€ into production operations.๐—•๐˜‚๐—ถ๐—น๐—ฑ ๐—ฎ ๐—ต๐—ถ๐—ด๐—ต-๐—ฝ๐—ฒ๐—ฟ๐—ณ๐—ผ๐—ฟ๐—บ๐—ถ๐—ป๐—ด ๐—œ๐—ป๐—ณ๐—ฟ๐—ฎ๐˜€๐˜๐—ฟ๐˜‚๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ, ๐—–๐—น๐—ผ๐˜‚๐—ฑ, ๐——๐—ฒ๐˜ƒ๐—ข๐—ฝ๐˜€ ๐—ฎ๐—ป๐—ฑ ๐—ฆ๐—ฅ๐—˜ ๐—ผ๐—ฟ๐—ด๐—ฎ๐—ป๐—ถ๐˜‡๐—ฎ๐˜๐—ถ๐—ผ๐—ป with clear accountability. 3. ๐—ž๐—ฒ๐˜† ๐—ฅ๐—ฒ๐˜€๐—ฝ๐—ผ๐—ป๐˜€๐—ถ๐—ฏ๐—ถ๐—น๐—ถ๐˜๐—ถ๐—ฒ๐˜€๐—–๐—น๐—ผ๐˜‚๐—ฑ & ๐—œ๐—ป๐—ณ๐—ฟ๐—ฎ๐˜€๐˜๐—ฟ๐˜‚๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ ๐—ฆ๐˜๐—ฟ๐—ฎ๐˜๐—ฒ๐—ด๐˜†๐—ข๐˜„๐—ป ๐˜๐—ต๐—ฒ ๐—ฐ๐—น๐—ผ๐˜‚๐—ฑ ๐—ถ๐—ป๐—ณ๐—ฟ๐—ฎ๐˜€๐˜๐—ฟ๐˜‚๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ ๐—ฟ๐—ผ๐—ฎ๐—ฑ๐—บ๐—ฎ๐—ฝ ๐—ฎ๐—ฐ๐—ฟ๐—ผ๐˜€๐˜€ ๐—”๐—ช๐—ฆ, ๐—”๐˜‡๐˜‚๐—ฟ๐—ฒ ๐—ฎ๐—ป๐—ฑ/๐—ผ๐—ฟ ๐—š๐—–๐—ฃ.Define scalable, secure, highly available and cost-effective cloud architecture and standards.Lead infrastructure modernization, technical-debt reduction, automation and multi-region/HA initiatives.Establish Infrastructure as Code and configuration-management standards. ๐—ฃ๐—ฟ๐—ผ๐—ฑ๐˜‚๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐—ข๐˜„๐—ป๐—ฒ๐—ฟ๐˜€๐—ต๐—ถ๐—ฝ & ๐—ฅ๐—ฒ๐—น๐—ถ๐—ฎ๐—ฏ๐—ถ๐—น๐—ถ๐˜๐˜†ยท Own end-to-end production health across customer environments, including Kubernetes, databases, storage, certificates, backups, capacity and critical dependencies.Establish proactive SRE practices, production readiness standards, operational ownership and runbooks.Ensure production environments are observable, supportable, secure and continuously improved.Drive availability, reliability and performance improvements for business-critical services.๐—œ๐—ป๐—ฐ๐—ถ๐—ฑ๐—ฒ๐—ป๐˜, ๐—ฃ๐—ฟ๐—ผ๐—ฏ๐—น๐—ฒ๐—บ & ๐—–๐—ต๐—ฎ๐—ป๐—ด๐—ฒ ๐— ๐—ฎ๐—ป๐—ฎ๐—ด๐—ฒ๐—บ๐—ฒ๐—ป๐˜ยท Establish ITIL-aligned incident, problem, change and escalation processes.ยท Lead major incidents and ensure effective on-call, severity and escalation models.ยท Own MTTA, MTTR, incident recurrence and post-incident improvement.ยท Ensure root-cause analysis is completed, and corrective actions are tracked to closure.ยท Strengthen change and release governance to reduce production risk.๐—ข๐—ฏ๐˜€๐—ฒ๐—ฟ๐˜ƒ๐—ฎ๐—ฏ๐—ถ๐—น๐—ถ๐˜๐˜† & ๐—ข๐—ฝ๐—ฒ๐—ฟ๐—ฎ๐˜๐—ถ๐—ผ๐—ป๐—ฎ๐—น ๐—œ๐—ป๐˜๐—ฒ๐—น๐—น๐—ถ๐—ด๐—ฒ๐—ป๐—ฐ๐—ฒยท Own monitoring, logging, metrics, tracing, alerting and operational dashboards.ยท Establish meaningful SLAs, SLOs and SLIs and use them to drive reliability decisions.ยท Reduce alert fatigue and improve proactive detection of performance, availability and capacity issues. ๐——๐—ฒ๐˜ƒ๐—ข๐—ฝ๐˜€, ๐—–๐—œ/๐—–๐—— & ๐—”๐˜‚๐˜๐—ผ๐—บ๐—ฎ๐˜๐—ถ๐—ผ๐—ปยท Drive CI/CD, Terraform/IaC, GitOps and automated deployment practices.ยท Promote reliable deployment strategies including zero-downtime, blue/green and canary releases where appropriate.ยท Automate repetitive operational tasks and eliminate avoidable manual intervention.ยท Partner with R&D to improve release readiness, deployment quality and production stability. ๐—ฅ๐—ฒ๐˜€๐—ถ๐—น๐—ถ๐—ฒ๐—ป๐—ฐ๐—ฒ, ๐——๐—ฅ & ๐—•๐˜‚๐˜€๐—ถ๐—ป๐—ฒ๐˜€๐˜€ ๐—–๐—ผ๐—ป๐˜๐—ถ๐—ป๐˜‚๐—ถ๐˜๐˜†ยท Design and maintain HA, backup, replication and failover strategies.ยท Define and validate RTO/RPO requirements.ยท Conduct regular backup/restore tests, failover exercises and DR simulations.ยท Ensure recoverability is demonstrated, documented and continuously improved. ๐—ฆ๐—ฒ๐—ฐ๐˜‚๐—ฟ๐—ถ๐˜๐˜†, ๐—ฅ๐—ถ๐˜€๐—ธ & ๐—–๐—ผ๐—บ๐—ฝ๐—น๐—ถ๐—ฎ๐—ป๐—ฐ๐—ฒยท Embed secure-by-design practices across cloud and production environments.ยท Govern IAM, least privilege, network segmentation, secrets, certificates, patching and vulnerability remediation.ยท Maintain security logging, hardening and operational controls.ยท Support audits, penetration testing, customer security assessments and regulatory/compliance requirements. ๐—–๐—ฎ๐—ฝ๐—ฎ๐—ฐ๐—ถ๐˜๐˜†, ๐—ฃ๐—ฒ๐—ฟ๐—ณ๐—ผ๐—ฟ๐—บ๐—ฎ๐—ป๐—ฐ๐—ฒ & ๐—–๐—น๐—ผ๐˜‚๐—ฑ ๐—–๐—ผ๐˜€๐˜ยท Establish proactive capacity planning, forecasting and performance management.ยท Identify infrastructure and application bottlenecks and coordinate improvements with R&D.ยท Drive cloud cost optimization and FinOps discipline without compromising reliability or security. ๐—Ÿ๐—ฒ๐—ฎ๐—ฑ๐—ฒ๐—ฟ๐˜€๐—ต๐—ถ๐—ฝ, ๐—ข๐—ฟ๐—ด๐—ฎ๐—ป๐—ถ๐˜€๐—ฎ๐˜๐—ถ๐—ผ๐—ป & ๐—ง๐—ฎ๐—น๐—ฒ๐—ป๐˜ยท Assess and evolve the Infrastructure/Cloud/DevOps/SRE operating model, roles and accountability.ยท Build a capable, scalable team with clear ownership, on-call and escalation structures.ยท Recruit, mentor and develop technical talent and strengthen operational discipline.ยท Provide hands-on technical leadership when complex production issues require intervention.๐—ฅ&๐—— & ๐—–๐—ฟ๐—ผ๐˜€๐˜€-๐—™๐˜‚๐—ป๐—ฐ๐˜๐—ถ๐—ผ๐—ป๐—ฎ๐—น ๐—ฃ๐—ฎ๐—ฟ๐˜๐—ป๐—ฒ๐—ฟ๐˜€๐—ต๐—ถ๐—ฝEstablish production-readiness requirements for new products, releases and customer environments.Lead cross-functional investigations across infrastructure, networking, Kubernetes, databases, storage, security and application architecture.Ensure production engineering is treated as a reliability and engineering functionโ€”not simply a ticket-routing/support function. ๐—ž๐—ป๐—ผ๐˜„๐—น๐—ฒ๐—ฑ๐—ด๐—ฒ, ๐——๐—ผ๐—ฐ๐˜‚๐—บ๐—ฒ๐—ป๐˜๐—ฎ๐˜๐—ถ๐—ผ๐—ป & ๐—š๐—ผ๐˜ƒ๐—ฒ๐—ฟ๐—ป๐—ฎ๐—ป๐—ฐ๐—ฒMaintain accurate architecture diagrams, infrastructure inventories, dependency maps and operational documentation.Establish and maintain runbooks for incidents, deployments, escalation, backup/restore and DR.Reduce dependency on individual knowledge through structured documentation and knowledge transfer.Ensure operational decisions and controls are traceable and audit-ready. ๐—˜๐˜…๐—ฒ๐—ฐ๐˜‚๐˜๐—ถ๐˜ƒ๐—ฒ ๐—ฅ๐—ฒ๐—ฝ๐—ผ๐—ฟ๐˜๐—ถ๐—ป๐—ด & ๐—ž๐—ฃ๐—œ๐˜€ยท Provide executive-level visibility on:ยท Availability, SLA/SLO performance and reliability trends.ยท Critical incidents, MTTA, MTTR and recurring issues.ยท Deployment/change success and production stability.ยท Capacity, cloud cost and operational efficiency.ยท Backup/restore and DR readiness.ยท Vulnerabilities, critical risks and technical debt.ยท Post-mortem actions and overall production health. 4. ๐—ง๐—ฒ๐—ฐ๐—ต๐—ป๐—ถ๐—ฐ๐—ฎ๐—น ๐—˜๐—ป๐˜ƒ๐—ถ๐—ฟ๐—ผ๐—ป๐—บ๐—ฒ๐—ป๐˜ & ๐—˜๐˜…๐—ฝ๐—ฒ๐—ฐ๐˜๐—ฒ๐—ฑ ๐——๐—ฒ๐—ฝ๐˜๐—ตยท Strong practical knowledge across:ยท AWS, Azure and/or GCP; multi-cloud/multi-region environments.ยท Kubernetes/OpenShift, Docker, Helm and ArgoCD/GitOps.ยท CI/CD, Terraform and Infrastructure as Code.ยท Linux, networking, DNS, TLS, reverse proxies and load balancing.ยท SQL databases, S3/object storage, Redis and Kafka.ยท Prometheus, Grafana, centralized logging, APM and distributed tracing.ยท Backup, restore, DR, HA and resilience engineering.ยท IAM, OpenID, LDAP/AD and cloud security controls.ยท APIs, distributed systems and application performance. 5. ๐—–๐—ฎ๐—ป๐—ฑ๐—ถ๐—ฑ๐—ฎ๐˜๐—ฒ ๐—ฃ๐—ฟ๐—ผ๐—ณ๐—ถ๐—น๐—ฒ & ๐—ค๐˜‚๐—ฎ๐—น๐—ถ๐—ณ๐—ถ๐—ฐ๐—ฎ๐˜๐—ถ๐—ผ๐—ป๐˜€12+ years of experience across IT Infrastructure, Cloud, DevOps, SRE or Production Operations.5+ years in senior technology leadership with ownership of critical, customer-facing production environments.Proven experience leading teams through major incidents, operational transformation and production governance.Strong understanding of SRE, cloud operating models, CI/CD, automation and ITIL-aligned practices.Demonstrated ability to balance executive-level strategy with hands-on technical intervention when required.Strategic, structured, proactive and highly accountable, with strong risk and operational discipline.Bachelorโ€™s/Masterโ€™s degree in Computer Science, Engineering, IT or a related field; MBA/Technology Management is an advantage. 6. ๐—ฃ๐—ฟ๐—ฒ๐—ณ๐—ฒ๐—ฟ๐—ฟ๐—ฒ๐—ฑ ๐—–๐—ฒ๐—ฟ๐˜๐—ถ๐—ณ๐—ถ๐—ฐ๐—ฎ๐˜๐—ถ๐—ผ๐—ป๐˜€AWS Solutions Architect Professional, Google Professional Cloud Architect, Azure Solutions Architect Expert, CKA, Terraform Associate, advanced SRE/DevOps certifications, CISSP, CISM and/or ITIL Managing Professional.