Senior Solutions Architect
Rafay(Asia) Pte. Ltd. · Singapore
Apply & track with Apply EdgeSenior Solutions Architect – Post-SalesFull Time — APAC (Singapore)About the RoleWe are seeking a Senior Solutions Architect to help customers successfully deploy, operate, and scale AI/ML workloads on our GPU Platform-as-a-Service (PaaS) offering. In this customer-facing role, you will work closely with platform engineering, MLOps, data science, and infrastructure teams to design and implement production-ready AI infrastructure solutions built on Kubernetes and GPU-accelerated environments.As a Senior Solutions Architect, you will serve as a trusted advisor to enterprise customers, lead complex deployments, drive architecture best practices, and help customers maximize the value of their AI infrastructure investments. You will help customers onboard to the platform, optimize workload performance, automate infrastructure, and ensure reliable operations while serving as a strategic technical partner throughout the customer lifecycle.This role is based in Singapore and covers the APAC region*Fluency in Mandarin is a key requirement for this position.ResponsibilitiesDesign end-to-end AI/ML platform architectures spanning inference, training, and data pipelinesDevelop reference architectures for GPU cluster deployment, LLM serving, and multi-tenant ML infrastructureEvaluate and recommend inference serving frameworks (vLLM, TGI, Triton, NIM)Advise on GPU fabric topology — NVLink, InfiniBand, RoCEv2 — for distributed trainingDesign observability strategies across DCGM, OTel, eBPF, and GPU metrics pipelinesTranslate complex infrastructure requirements into actionable platform designsDeliver technical presentations, workshops, and proof-of-concept engagementsAct as a trusted advisor on AI infrastructure strategy, cost, and scalingPartner with customer platform, MLOps, data science, and executive stakeholders to understand AI/ML workload requirements and translate them into scalable platform architectures.Architect networking, identity management, observability, and security integrations with enterprise systems.Monitor and troubleshoot production environments, including GPU utilization, workload performance, cluster health, and cost efficiency.Lead root cause analysis and remediation efforts for complex customer issues.Serve as the primary technical advisor and escalation point for assigned customers across the APAC region.Provide technical leadership during customer engagements and drive adoption of platform best practices.Document reference architectures, implementation guides, and best practices.Provide feedback to Product and Engineering teams to improve platform capabilities and influence product roadmap decisions.Collaborate with internal teams to ensure successful customer adoption, expansion, and long-term success.Mentor junior team members and contribute to the growth of the Solutions Architecture organization.Required Qualifications8+ years in infrastructure, platform, or solutions engineering roles3+ years focused on AI/ML infrastructure or MLOpsDeep Kubernetes expertise — cluster lifecycle, workloads, operators, RBACHands-on experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred)Proficiency with distributed training concepts — NCCL, tensor parallelism, pipeline parallelismExperience with LLM inference serving and optimization (vLLM, NIM, TGI)Familiarity with GPU Operator, MIG, SR-IOV, and network fabrics (IB/RoCEv2)Strong scripting and automation skills (Python, Bash, Go preferred)Demonstrated ability to communicate complex technical concepts to diverse audiencesExperience with at least one programming language such as Python or Go.Experience with AWS, Azure, or GCP, including networking, IAM, and managed Kubernetes services.Familiarity with monitoring and observability technologies including Prometheus, Grafana, OpenTelemetry, or similar.Strong understanding of AI/ML infrastructure concepts including GPU-based workloads, model serving, training pipelines, and resource optimization.Proven ability to troubleshoot and resolve complex infrastructure and platform issues.Excellent communication, presentation, and customer-facing skills.Experience leading technical discussions with both engineering teams and executive stakeholders.Fluent in Mandarin and English (spoken and written) Willingness to travel regularly across the APAC region.Preferred QualificationsExperience supporting enterprise customers in cloud-native environments.Familiarity with AI/ML frameworks such as PyTorch and TensorFlow.Experience with Run:AI, Slurm,Experience with GPU scheduling, autoscaling, and workload optimization.Understanding of multi-tenant Kubernetes environments and platform operations.Experience working with MLOps or AI infrastructure platforms.Experience developing reference architectures and leading technical workshops.Relevant certifications such as CKA, CKAD, AWS Solutions Architect, Azure Solutions Architect, or GCP Professional Cloud Architect.Understanding of multi-tenant GPU isolation (SR-IOV VFs, DPU offload)Experience covering Southeast Asia and/or Greater China enterprise accounts.