أبلاي إيدج ابدأ البحث عن عمل

AI Platform Engineer

Sharon AI, Inc · United Arab Emirates

قدّم وتابع مع أبلاي إيدج
AI Platform Engineer L3 (GPUaaS – AI Neocloud)📍 EMEA | Remote-firstAbout Sharon AISharon AI is building the infrastructure powering the next generation of artificial intelligence.Operating across AI infrastructure, high-performance compute, cloud platforms and large-scale technology environments, Sharon AI delivers scalable, secure and reliable infrastructure for demanding AI, ML and HPC workloads.The RoleAs an AI Platform Engineer L3, you'll design, build and operate the platform layer powering Sharon AI's GPU-as-a-Service (GPUaaS) offering across the EMEA region. You'll own the architecture, automation and reliability of the platform services sitting above Sharon AI's GPU and network fabric, spanning Kubernetes, Slurm, GPU scheduling, MLOps tooling, model serving and platform observability across multiple EMEA sites.Reporting to the Head of Operations, you'll work closely with Network Engineering, Infrastructure and customer-facing teams to solve complex, cross-team challenges and ensure Sharon AI's platform can scale reliably and efficiently across the region. This is a hands-on senior individual contributor role suited to an engineer with strong platform engineering, DevOps or MLOps experience who can operate independently in a fast-paced, AI-native environment and provide technical guidance to less experienced engineers.Key ResponsibilitiesDesign and own the AI platform architecture across Sharon AI's EMEA GPU clusters, including Kubernetes, Slurm and container orchestrationLead the development of CI/CD pipelines and MLOps tooling supporting training, fine-tuning and inference workloads across multiple sitesDefine and implement multi-tenant GPU resource scheduling, quota management and workload isolation strategies at scaleOwn the design of model serving infrastructure, balancing high availability, performance and cost efficiencyBuild and evolve platform-wide observability across monitoring, logging and alerting, covering platform health, GPU utilisation and workload performanceDrive Infrastructure-as-Code adoption and platform automation using Terraform and AnsiblePartner closely with Network Engineering to integrate the platform layer with Sharon AI's underlying InfiniBand/RDMA fabricAct as a senior escalation point for complex platform issues impacting customer AI/ML workloads across EMEALead incident response and post-incident reviews, contributing to operational runbooks and platform best practicePartner with the Head of Operations on platform capacity planning, scaling strategy and cost optimisation across EMEAMentor and provide technical guidance to less experienced platform engineersSupport enterprise GPUaaS customer onboarding and technical escalations across the regionSkills & Experience6–10+ years' experience in platform engineering, DevOps, MLOps or SRE, ideally within HPC, cloud or AI/ML infrastructure environmentsBachelor's degree in Computer Science, Electrical Engineering or a related fieldHands-on experience operating Kubernetes and GPU scheduling at production scaleProven experience designing CI/CD and Infrastructure-as-Code practices for platform teamsProven experience supporting GPU or AI/ML infrastructure at scale, ideally within a GPUaaS or neocloud environmentDeep expertise in Kubernetes and GPU scheduling frameworks, including Slurm, Kubernetes device plugins and NVIDIA GPU OperatorStrong experience designing and operating MLOps tooling and ML pipeline orchestration in productionAdvanced proficiency in Python, Bash and Infrastructure-as-Code tools such as Terraform and AnsibleStrong understanding of GPU infrastructure and distributed training concepts, including NCCL, data/model parallelism and RDMA-aware schedulingExperience with platform-wide observability tooling such as Prometheus, Grafana and telemetry stacksStrong Linux systems and networking fundamentals, with the judgement to independently solve ambiguous, cross-team problemsA security-first mindset when operating within multi-tenant environmentsAwareness of EMEA regulatory and data residency considerations, including GDPR, as they relate to platform operationsStrong communication and collaboration skills across distributed, multi-region teams, with the ability to mentor other engineersExperience with distributed training frameworks such as PyTorch and TensorFlow, MLOps platforms including MLflow, Kubeflow or Ray, and InfiniBand/RDMA or RoCEv2 networking concepts is advantageous. Kubernetes certifications such as CKA, CKAD or CKS, and cloud certifications across AWS, GCP or Azure, are also advantageous. The role requires the right to work in an EMEA jurisdiction, with existing eligibility to work across the EU/EEA or UK advantageous given the multi-country remit.Why Join Sharon AIOwn the platform architecture powering a growing GPU-as-a-Service and AI neocloud business across EMEAWork hands-on with large-scale GPU infrastructure, Kubernetes, Slurm and AI-native platform technologiesShape how Sharon AI's platform scales across multiple jurisdictions, sites and customer workloadsSolve complex technical challenges across GPU infrastructure, MLOps, networking and distributed AI workloadsInfluence platform reliability, automation, capacity and cost optimisation across the regionWork closely with Network Engineering and Infrastructure teams on high-performance InfiniBand/RDMA environmentsProvide technical leadership and mentorship while remaining hands-on as a senior individual contributorHelp enterprise customers reliably train, fine-tune and run AI/ML workloads at scaleJoin a highly technical and ambitious team operating at the forefront of AI infrastructureOur ValuesIntegrity | Innovation | Collaboration | Wellbeing | Inclusion