أبلاي إيدج ابدأ البحث عن عمل

Senior Site Reliability Engineer (Azure)

Salt · Abu Dhabi, Abu Dhabi Emirate, United Arab Emirates

قدّم وتابع مع أبلاي إيدج
Location: Abu Dhabi, UAE
Employment: 12-months initial (contract to perm)We are seeking a hands-on Senior DevOps / Site Reliability Engineer to build and operate the delivery and runtime foundations for a portfolio of modern enterprise applications, workflow platforms and AI-enabled products.This is not primarily an infrastructure-administration role. You will work directly with software engineers to create reliable, secure and automated paths from code to production.You will own deployment automation, runtime reliability, observability, infrastructure-as-code and operational readiness across applications that integrate with critical enterprise systems.What You Will OwnPlatform & Infrastructure* Build and maintain cloud infrastructure using Infrastructure as Code.* Design secure, repeatable environments across development, test, staging and production.* Manage containerized workloads and Kubernetes-based deployments where appropriate.* Define standard application deployment patterns for backend, frontend and AI services.* Implement secure secrets and configuration management.* Support network, identity and connectivity requirements for enterprise integrations.CI/CD & Developer Productivity* Build automated CI/CD pipelines.* Standardize build, test, security scanning and deployment processes.* Automate environment provisioning and configuration.* Reduce manual deployment steps and production configuration drift.* Work closely with engineering teams to improve release frequency and reliability.Reliability & Observability* Establish logging, metrics, tracing and alerting.* Define service-level indicators and operational thresholds.* Build dashboards for system health and application performance.* Implement incident-response and production-support practices.* Design for graceful degradation, retries, failover and recovery.* Lead root-cause analysis of production incidents.Security & Operational Controls* Implement least-privilege access and secure deployment patterns.* Support auditability of infrastructure and production changes.* Integrate security checks into delivery pipelines.* Work with security and infrastructure teams to meet enterprise control requirements.Resilience* Support business continuity and disaster-recovery design.* Define backup, restore and recovery procedures.* Test operational recovery rather than relying solely on documented plans.Required Experience* 6+ years in DevOps, SRE, platform engineering or cloud infrastructure.* Strong production experience with Azure, AWS or GCP; Azure strongly preferred.* Docker and Kubernetes.* Infrastructure as Code using Terraform, Bicep, Pulumi or equivalent.* CI/CD using Azure DevOps, GitHub Actions, GitLab CI or similar.* Strong Linux and networking fundamentals.* Observability tooling and distributed-system troubleshooting.* Secure secrets, identity and access-management patterns.* Production incident-management experience.* Scripting/programming capability in Python, Go, Bash or equivalent.Strong Advantage* Azure Kubernetes Service.* Azure Service Bus, API Management, Key Vault and related Azure services.* Enterprise integration platforms.* SAP-connected environments.* AI/LLM application deployment.* Regulated or government environments.* High-availability and disaster-recovery architecture.