Apply Edge Start your job search

Senior DevOps Engineer – Kubernetes

Socium - Teams Done Differently · Dubai, Dubai, United Arab Emirates

Apply & track with Apply Edge

General DescriptionWe are looking for an experienced Senior Site Reliability Engineer (SRE) / Kubernetes Infrastructure Engineer with strong hands-on expertise across on-premise Kubernetes, Linux infrastructure, air-gapped environments, infrastructure migrations, cluster troubleshooting, PostgreSQL, distributed storage, container registries, and production operations.The successful candidate will be responsible for deploying, upgrading, migrating, troubleshooting, and supporting on-premise Kubernetes environments within highly secure client data centers, including air-gapped and restricted-access environments where internet connectivity, personal devices, and remote technical assistance may not be available.This is a highly hands-on, client-facing infrastructure role requiring someone who can operate independently and confidently within production environments. Candidates must be comfortable taking ownership of complex infrastructure changes, diagnosing failures directly from the Linux command line, making technical decisions under pressure, and clearly communicating technical status and remediation plans to client stakeholders.The role requires deeper Kubernetes expertise than standard cloud-managed Kubernetes administration. Candidates should have practical experience with Kubernetes control plane and data plane components, etcd, kubeadm, certificate management, cluster networking, PostgreSQL replication, private container registries, and Linux infrastructure.Experience supporting government, cybersecurity, enterprise infrastructure, systems integration, professional services, or other high-security environments would be highly advantageous.The position supports an AI / GenAI cybersecurity product, providing security capabilities for enterprise AI environments, making this an excellent opportunity for infrastructure engineers interested in working at the intersection of Kubernetes, infrastructure engineering, cybersecurity, and AI platforms.Key ResponsibilitiesOwn the deployment and operational support of on-premise Kubernetes environments across enterprise and government client infrastructurePlan and execute Kubernetes cluster deployments, upgrades, migrations, and reconfigurationsPerform complex infrastructure changes including IP re-addressing, certificate rotation, cluster reconfiguration, and production migrationsDeploy and support Kubernetes within air-gapped, restricted-access, and high-security data center environmentsTroubleshoot complex Kubernetes and infrastructure failures independently while working directly at client sitesManage and troubleshoot both Kubernetes control plane and data plane componentsDiagnose issues involving etcd quorum, peer membership, cluster state, and failure recoveryBuild and maintain Kubernetes clusters using kubeadmManage Kubernetes certificates including SAN configuration, certificate renewal, CA rotation, and certificate troubleshootingSupport PostgreSQL primary/replica replication and troubleshoot database infrastructure issuesManage and troubleshoot Linux networking, including DNS, NTP, firewall rules, routing, connectivity, and related infrastructure dependenciesSupport private container registries such as Docker Distribution or equivalent technologiesWork with distributed storage platforms such as SeaweedFS, Ceph, MinIO, or similar technologiesSupport centralized logging environments using ELK or equivalent logging and observability platformsWork extensively through the Linux command line to diagnose and resolve infrastructure issuesWrite, review, and troubleshoot Bash scripts, deployment scripts, installers, and infrastructure automation toolingValidate deployment scripts, installers, automation, and infrastructure changes thoroughly within lab and staging environments before production executionRepresent technical infrastructure activities directly to client stakeholders during deployments, migrations, incidents, and go-live activitiesClearly communicate technical failures, root causes, risks, remediation plans, and deployment statusProduce clear and structured runbooks, operational procedures, decision trees, deployment documentation, and incident reportsIdentify and escalate infrastructure and deployment risks proactivelyTravel to client data centers across the UAE as required for deployment, migration, troubleshooting, and go-live supportOperate independently within environments where remote assistance or internet connectivity may not be availableSupport infrastructure requirements for secure AI / GenAI cybersecurity platformsWhere applicable, support GPU-enabled Kubernetes infrastructure, including NVIDIA device plugins and container tooling.Required ExperienceApproximately 5–6+ years of strong hands-on Kubernetes experience within production environmentsProven professional experience deploying or operating at least one on-premise Kubernetes environmentCandidates whose experience is exclusively with managed cloud Kubernetes services such as AKS, EKS, or GKE without on-premise Kubernetes exposure may not be suitableStrong hands-on understanding of Kubernetes architecture, control plane, data plane, cluster networking, scheduling, storage, and production operationsStrong understanding of etcd internals, including quorum, peer membership, cluster health, failure scenarios, and recoveryHands-on experience with kubeadm-based Kubernetes cluster bootstrappingPractical experience managing Kubernetes certificates, SANs, CA rotation, renewal, and certificate-related troubleshootingStrong Linux administration and troubleshooting experienceComfortable operating extensively from the Linux command lineAbility to write, review, troubleshoot, and modify Bash scriptsStrong understanding of Linux networking fundamentals including DNS, NTP, firewalls, routing, ports, and connectivity troubleshootingWorking knowledge of PostgreSQL, ideally including primary/replica replicationExperience working with private container registries, Docker Distribution, or equivalent technologiesExperience troubleshooting infrastructure issues directly within production environmentsDemonstrated ability to work independently without relying continuously on remote engineering supportExperience validating infrastructure changes within lab, test, or staging environments before production deploymentStrong incident management and troubleshooting capabilitiesStrong client-facing communication skills with the ability to explain technical issues clearly to both technical and non-technical stakeholdersAbility to remain structured and effective within high-pressure and high-stakes production environmentsComfortable travelling to client sites and secure data centers as requiredComfortable working within air-gapped or restricted-access environmentsMust be available to work onsite in the UAEPreferred ExperienceHands-on experience working with air-gapped Kubernetes environmentsExperience supporting Kubernetes within government, defense, cybersecurity, banking, critical infrastructure, or other highly secure environmentsExperience with distributed storage technologies such as SeaweedFS, Ceph, MinIO, or equivalentExperience with centralized logging and observability technologies such as ELK / Elasticsearch / Logstash / KibanaExperience with GPU-enabled Kubernetes nodesExposure to NVIDIA Device Plugin, NVIDIA Container Toolkit, GPU scheduling, or AI/ML infrastructurePrevious experience supporting infrastructure for AI, Machine Learning, Generative AI, or cybersecurity platformsPrevious consulting, systems integration, professional services, or client-facing infrastructure engineering experienceExperience working directly with enterprise or government clientsPrevious professional experience within the UAE or wider GCC / Middle East regionUnderstanding of infrastructure security concepts including access controls, credential handling, certificate security, network segmentation, and secure operational practicesExperience creating detailed runbooks, deployment procedures, troubleshooting guides, incident reports, and operational documentationExperience executing infrastructure migrations and production changes where remote support or internet connectivity is unavailable.Work Setup