Apply Edge Start your job search

AI DevOps Engineer (Sovereign Infrastructure)

Xpin AI · Dubai, United Arab Emirates

Apply & track with Apply Edge
Own the infrastructure that puts our AI into environments most platforms never reach: air-gapped government data centres running GPU clusters with no route to the internet, alongside cloud on Azure and GCP.About Xpin AIXpin AI builds sovereign AI for governments, enterprises, and institutions across the GCC. Our products are bilingual at their core, spanning Arabic, English, and regional languages, and ship into both on-premise sovereign environments and cloud. The infrastructure problem is the hard one: the same product runs on a customer's own GPU hardware behind an air gap and on managed cloud, without two divergent codebases.The roleYou are the person who makes deployment repeatable. Reaching a sovereign customer environment today takes too much manual work and too much institutional memory. You will turn that into infrastructure as code, GitOps, and an offline artifact pipeline another engineer can run without you in the room.You will work in a squad alongside AI and product engineers, and spend real time at client sites standing up GPU clusters and debugging problems you cannot reproduce from your desk. This is a senior individual contributor role with architectural ownership. Within a year, a new sovereign environment should go from bare hardware to running product in days rather than weeks, with observability and rollback that work the same on a customer data centre or Azure.What you'll buildThe offline software supply chain. Mirrored package registries, a private container registry, and and a signed artifact promotion path that moves builds across the air gap auditably. This is what decides whether a sovereign deployment ships on time.GPU cluster provisioning on bare metal. Stand up and operate on-premise clusters on customer hardware including NVIDIA H200 servers: driver and CUDA lifecycle, GPU scheduling and partitioning, and capacity planning against hardware you do not choose.One deployment path across sovereign and cloud. Infrastructure as code and GitOps targeting an air-gapped cluster, Azure, or GCP from one declarative source, so environments do not drift into snowflakes.Cloud-native operations on Azure. Run and harden managed Kubernetes: identity, private networking, policy and cost governance, so cloud deployments meet the same security bar as sovereign ones.Model and inference serving. GPU-aware scheduling, autoscaling where hardware permits, model weight distribution and versioning, and the tuning that holds latency inside customer expectations.The model lifecycle pipeline. Validation gates, staged rollout, versioning for models, datasets, and experiments, and drift monitoring once live. A customer may ask which model and data produced a result.Observability with no egress. A self-hosted metrics, logs, and traces stack including GPU telemetry, with nothing leaving the network.Security and compliance as a design constraint. Hardened OS and cluster baselines, policy enforcement, secrets management, image scanning and SBOM generation, and the evidence trail security reviews ask for.Backup, disaster recovery, and runbooks. Tested restore paths and documentation good enough that an on-call engineer new to the environment can act at 3am.Tech stackKubernetes and orchestration: RKE2 and k3s on-premise, AKS on Azure, GKE on GCP; Helm and Kustomize; Cilium for networking and network policy.GPU and hardware: NVIDIA H200 servers, NVIDIA GPU Operator, CUDA and driver lifecycle, MIG partitioning, DCGM telemetry, NVLink and RDMA or InfiniBand fabrics, PXE and out-of-band provisioning (iDRAC, IPMI, Redfish).Infrastructure as code and config: Terraform, Azure Bicep, Ansible for bare metal, Packer for image builds.GitOps and CI/CD: Argo CD or Flux, GitHub Actions with self-hosted runners, self-hosted GitLab CI or Gitea where the environment has no route out, air-gapped pipeline patterns.Artifacts and supply chain: Harbor, MinIO, Nexus or Artifactory for mirrored package registries, Cosign and Sigstore for image signing, Syft and Grype or Trivy for SBOM and vulnerability scanning.Cloud, Azure first: AKS, Azure Arc for hybrid and sovereign cluster management, Entra ID, Key Vault, Azure Container Registry, Azure Policy, Private Link, Azure Monitor. GCP second: GKE, Artifact Registry, Workload Identity, Cloud Armor.Observability: Prometheus, Grafana, Loki, Tempo, OpenTelemetry, Alertmanager, all self-hosted.Security and policy: HashiCorp Vault or External Secrets, SOPS or Sealed Secrets, Kyverno or OPA Gatekeeper, Falco, CIS benchmark hardening.Storage and data: Ceph, Longhorn, or NFS on-premise; PostgreSQL operators; Velero for backup and restore.Serving and workloads: vLLM and Triton Inference Server deployments, KEDA for event-driven scaling.MLOps: self-hosted MLflow or Weights & Biases for experiment tracking and model registry, DVC or LakeFS for dataset versioning, Argo Workflows or Kubeflow for pipeline orchestration, drift and performance monitoring.Traffic and mesh: NGINX or Traefik ingress, Istio or Linkerd where a mesh is warranted, mTLS between services.Languages and tooling: Python, including enough FastAPI to read and modify the services you deploy; Bash; Go where the tooling calls for it; Linux administration at depth.Level and scopeSenior, roughly five to eight years in DevOps, platform, or SRE work as a guide rather than a gate, with real time on both bare metal and managed cloud rather than only one. You own the deployment and operations layer across your projects, set the standards others follow, and make architectural calls with the engineering leads. There is a path to platform lead as the team grows. Given how recent parts of this stack are, we weight demonstrated depth and strong Linux and Kubernetes fundamentals over exact tool tenure.Must-haveHas run Kubernetes in production and can debug it below the abstraction: networking, storage, scheduling, and the failure modes that never appear in a dashboard.Has run production workloads on a major cloud and owns the identity, networking, and cost decisions involved. Azure is our primary cloud and AKS is what we look for first, though strong GCP or AWS depth transfers well.Deep Linux and networking fundamentals, and comfort on bare metal as well as managed compute.Has built infrastructure as code and CI/CD pipelines other engineers actually use, and treats manual deployment steps as a defect.Has operated in restricted or no-egress environments, or can show clear evidence of understanding what breaks without internet access and how to design around it.Comfortable in Python at a level where you can read, debug, and modify the application code you deploy rather than treat it as a black box.Treats security and compliance as part of the build rather than a review gate at the end.Writes runbooks that survive their own absence, automates anything done twice, and communicates clearly with customer IT and security teams who hold real authority over your deployment.GPU infrastructure: NVIDIA drivers and operators, scheduling, MIG, or multi-node inference fabrics.Serving machine learning or LLM workloads in production, including inference tuning.Willing to travel to GCC client sites and work hands-on with physical hardware.Based in the UAE or willing to relocate to Dubai, and comfortable working on-site.Nice-to-haveMLOps at depth: automated retraining, model registries, dataset lineage, or production drift detection.Building backend services and APIs around model serving with FastAPI or similar.Working familiarity with PyTorch, LangChain, or LlamaIndex, enough to debug across the boundary with the AI engineering team.Advanced Azure beyond running clusters: Arc-enabled Kubernetes for estate outside Azure, landing zone design, or cost governance at scale.Supply chain security: image signing, SBOM generation, provenance, and pipeline-level vulnerability management.Delivery into government, defence, banking, or other regulated environments and their compliance frameworks.Ceph or other on-premise distributed storage at production scale.Certifications such as CKA, CKS, or Azure Solutions Architect, though ability counts for more.Arabic language ability, not required.Why joinSovereign AI infrastructure is a small field and this is a real example of it: air-gapped GPU clusters, frontier hardware, and a genuine hybrid problem across on-premise and two clouds. Most platform engineers never get to build any of it. You will build the platform rather than maintain someone else's, with the architectural decisions still open and your name on them. Expect a learning budget and room to grow into platform leadership.Compensation and logisticsOn-site in Dubai, with travel to client sites across the GCC. Compensation is competitive and benchmarked to your experience and the UAE market. The package includes UAE employment visa sponsorship, medical insurance, and annual leave and flight allowance per UAE labour law and company policy. Relocation support is available from outside the UAE. There is a shared on-call rotation for production environments. Xpin AI hires on demonstrated ability and welcomes applicants from all backgrounds.How to applySend your CV along with anything that shows how you work: a GitHub profile, Terraform or Ansible you have written, a homelab, or a technical blog. We care more about what you have run than what you have listed. A note on the hardest deployment you have shipped, and what went wrong, tells us more than a certification list.