Senior Inference Performance Engineer (AI)
GreenNode · Ho Chi Minh City, Vietnam
Apply & track with Apply EdgeJob Description:We are looking for a Senior Inference Engineer with a strong foundation in software engineering, distributed systems, and performance optimization to build and optimize inference engines for large-scale LLM serving systems. You will work across both research and production environments, ensuring our LLM serving systems are fast, scalable, and efficient. The role spans the entire inference stack — from kernel and runtime to scheduling, memory management, and distributed executionKey Responsibilities:Profile, benchmark, and analyze bottlenecks for LLM inference workloads across multiple layers: kernel, memory, networking, and schedulerOptimize inference engines (vLLM, SGLang, TensorRT-LLM) for throughput, latency, memory efficiency, GPU utilization, and costImplement and fine-tune inference optimization techniques including batching, KV-cache management, quantization, speculative decoding, parallelism strategies, and disaggregated servingBuild instrumentation and profiling tools to identify bottlenecksEnsure the reliability of the inference pipeline through A/B launches, rollback, model versioning, and fault toleranceCollaborate with the Platform Engineering team to improve serving architecture based on performance findingsDocument and share knowledge, contributing to internal best practices and AI open-source projects whenever possibleJob Requirements:Minimum 5 years of experience as a Software Engineer, Performance Engineer, or equivalent rolesStrong foundation in software engineering, distributed systems, and performance optimizationProficient in Python; experience with C/C++, Go, or other low-level programming languages is a plusExperience with high-performance systems such as high-throughput backends, distributed services, large-scale serving systems, or equivalentKnowledge of or hands-on experience with AI/ML serving systems, LLM inference, or GPU workloadsUnderstanding of GPU architecture and the CUDA programming modelHands-on experience with at least one inference engine (vLLM, SGLang, TensorRT-LLM, Triton Inference Server) or equivalent systemsUnderstanding of inference optimization techniques including batching, KV-cache optimization, quantization, speculative decoding, tensor/pipeline parallelism, or disaggregated servingExperience with profiling, bottleneck analysis, and performance tuning for production systemsStrong systems thinking, a high sense of ownership, and willingness to dive deep into complex technical challengesNice to Have:Hands-on production experience with CUDA programming or NVIDIA libraries such as cuBLAS, cuDNN, and NCCLOpen-source contributions to inference-related projects (vLLM, SGLang, TensorRT-LLM) or AI infrastructure projectsDeep experience with the NVIDIA inference stack including Triton, CUTLASS, TensorRT, or CUDA profiling tools (Nsight Systems / Nsight Compute)Experience with distributed inference, request routing, or inference orchestraExperience with LLMOps, RAG systems, or AI agtionentsExperience with observability stacks (Prometheus, Grafana, OpenTelemetry) for ML or distributed systemsPublished research papers or open-source contributions in the field of ML systems or inference optimization