أبلاي إيدج ابدأ البحث عن عمل

Platform Engineer (m/f/d)

Halian | Managed Services, Recruitment Agency & Contract Staffing · Abu Dhabi Emirate, United Arab Emirates

قدّم وتابع مع أبلاي إيدج
Platform Engineer6 month extendable contract Abu Dhabi | Office based We are looking for a highly skilled Platform Engineer / Site Reliability Engineer (SRE) to support and enhance our cloud-native platform powering customer-facing products. This role is heavily focused on Kubernetes operations, debugging, troubleshooting, performance optimization, and reliability engineering across production environments.You will work closely with Software Engineering, Product, Security, and Operations teams to ensure our applications and platform services remain highly available, scalable, secure, and observable. The ideal candidate combines deep Kubernetes expertise with strong application-layer troubleshooting skills and experience supporting complex distributed systems in a product-driven environment.This is a hands-on technical role requiring strong analytical and problem-solving skills, with the ability to diagnose issues across the full technology stack, from infrastructure and networking to application runtime and service dependencies.Key Responsibilities Own and support production Kubernetes environments, including cluster health, node management, networking, ingress, storage, and workload performance. Lead troubleshooting and root cause analysis of Kubernetes-related incidents impacting production systems. Debug complex issues involving pods, deployments, services, ingress controllers, networking, DNS, service meshes, and platform components. Optimize cluster performance, capacity planning, resource allocation, scaling, and resilience. Drive platform reliability through automation, monitoring, and proactive operational improvements. Act as an escalation point for critical production incidents and application outages. Troubleshoot issues across multiple layers including:Kubernetes Linux Containers Networking Databases APIs Microservices CI/CD pipelines Partner with development teams to diagnose application performance bottlenecks, deployment failures, memory leaks, connectivity issues, and latency concerns. Perform root cause analysis and implement permanent corrective actions following incidents. Improve application reliability, fault tolerance, and operational readiness. Define and maintain Service Level Objectives (SLOs), SLIs, and operational metrics. Enhance observability through logging, monitoring, tracing, and alerting. Create and maintain operational runbooks, dashboards, and incident response procedures. Participate in production support, major incident management, and on-call rotations. Drive automation to reduce manual operational effort and improve platform stability. Build and maintain Infrastructure as Code using Terraform. Support GitOps-based deployments and platform configuration management. Improve CI/CD pipelines and deployment reliability. Automate operational tasks, remediation workflows, and environment provisioning. Contribute to platform standards, architecture improvements, and engineering best practices. Implement Kubernetes and cloud security best practices. Collaborate with security teams on vulnerability remediation and compliance requirements. Support platform governance through policy enforcement, access controls, and audit readiness. Continuously improve platform resilience, recovery processes, and disaster recovery capabilities. Required Skills & Experience 5+ years of experience in Platform Engineering, Site Reliability Engineering, DevOps, or Cloud Infrastructure roles. Advanced hands-on experience supporting production Kubernetes environments (AKS, EKS, GKE, or OpenShift). Proven expertise in Kubernetes troubleshooting and debugging. Strong understanding of container technologies including Docker and Kubernetes internals. Experience troubleshooting distributed microservices-based applications. Strong Linux administration and system troubleshooting skills. Solid understanding of cloud platforms, preferably Azure. Experience with Infrastructure as Code using Terraform. Experience supporting CI/CD pipelines and deployment automation. Strong understanding of networking concepts including:DNS Load balancing Ingress controllers TLS/SSL Network policies Service-to-service communication Experience with monitoring, logging, and observability platforms. Strong incident management and root cause analysis experience. Excellent communication skills and ability to work directly with software engineering tea Platform Engineer in Abu Dhabi, United Arab Emirates