Senior Cloud DevOps Engineer – AI Cloud
Dada Consultants · Ayer Itam, Penang, Malaysia
Apply & track with Apply EdgeCloud Senior DevOps Engineer
They design and manufacture their own computing hardware, operate large-scale datacenters, and provide advanced AI cloud services to enterprise customers with demanding workloads. Headquartered in Singapore, the company is scaling rapidly and offers a rare opportunity to work on both Bitcoin mining infrastructure and cutting-edge AI/HPC systems at global scale. This role is central to the reliability, security, and continuous delivery of mission-critical platforms — making it an exceptional opportunity for a senior infrastructure engineer looking to operate at the frontier of cloud-native and AI infrastructure.Key ResponsibilitiesDesign, implement, and maintain end-to-end CI/CD and MLOps pipelines for software applications and machine learning models, automating build, test, deployment, and rollback processes to ensure smooth transitions from development to productionAct as the technical lead during complex system incidents — driving rapid troubleshooting, thorough root cause analysis, and the implementation of preventative remediation plansEstablish and enforce robust security and compliance standards, including release workflow governance, Zero Trust access controls, secrets management, and adherence to frameworks such as SOC2 and ISO27001Collaborate closely with R&D, Data Science, Security, and Business teams to eliminate workflow bottlenecks and drive engineering efficiency through Internal Developer Platform (IDP) and Platform Engineering initiativesArchitect and continuously improve comprehensive observability systems — including monitoring, logging, and alerting stacks (e.g., Prometheus, Grafana, ELK/EFK) — to maintain deep visibility into system health, application performance, and AI model metricsChampion Infrastructure as Code (IaC) practices using tools such as Terraform, Ansible, and Helm to achieve fully automated, reproducible, and auditable infrastructure provisioning across multi-cloud environmentsOwn high-availability architecture design in production, including disaster recovery strategies, self-healing mechanisms, capacity planning, and performance tuning to meet stringent business SLAsBuild, optimise, and scale cloud-native infrastructure using Kubernetes and Docker, including provisioning and managing GPU clusters to support high-performance AI workloads and model inferencingRequirementsBachelor's degree or above in Computer Science, Engineering, or a related technical discipline, with at least 5 years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure rolesSolid domain knowledge across CI/CD methodologies, Infrastructure as Code, observability paradigms, and SRE principles, combined with strong problem-solving ability and cross-functional communication skillsStrong scripting and programming proficiency in at least one major language such as Go, Python, or Shell, with a clear focus on automation and internal tooling developmentDemonstrated experience designing and managing infrastructure on major public or hybrid cloud platforms (e.g., AWS, GCP, Azure, Alibaba Cloud), including multi-cloud strategiesDeep hands-on expertise with Docker and Kubernetes orchestration, including cluster management and production-level best practices, alongside expert-level Linux and core networking knowledge (TCP/IP, DNS, HTTP, load balancing, VPCs)Experience with MLOps practices and model serving frameworks (e.g., vLLM, TGI, Triton Inference Server), and managing GPU clusters for AI/ML workloads (is a bonus)Prior experience in large-scale distributed or high-concurrency environments (e.g., fintech, real-time processing, or AI platforms), and/or experience building Internal Developer Platforms (IDP) (is a bonus)Experience in a technical lead or mentoring capacity within a DevOps or SRE team (is a bonus)