AI Runtime Engineer
XpertDirect · Vienna, Austria
Apply & track with Apply EdgeAI Runtime EngineerVienna, Austria — HybridAI Infrastructure | Distributed Systems | Large Language Models | High-Performance ComputingOur client, an innovative AI Infrastructure company based in Vienna, is building the runtime platform that powers high-performance AI applications used by enterprise customers across Europe.They're looking for an AI Runtime Engineer to help optimise the systems responsible for serving Large Language Models across distributed GPU infrastructure, where every millisecond of latency and every percentage of GPU utilisation matters.This is a deeply technical engineering role sitting at the intersection of distributed systems, cloud infrastructure, and modern AI.Your ResponsibilitiesDevelop high-performance runtime services responsible for serving production AI modelsOptimise inference performance across distributed GPU infrastructureBuild scalable model serving systems capable of handling enterprise AI workloadsDevelop platform services and performance tooling using Rust and PythonImprove scheduling, orchestration, and resource utilisation across Kubernetes clustersWork closely with AI Researchers and Machine Learning Engineers to optimise production deploymentProfile system performance and remove bottlenecks across networking, memory, and compute layersImprove platform observability, reliability, and operational efficiencyContribute to the architecture of the company's next-generation AI infrastructure platformExperience Required5+ years of experience in Backend Engineering, Platform Engineering, Distributed Systems, or AI InfrastructureStrong commercial experience developing software in Rust or modern systems programming languagesExcellent Python development skillsHands-on experience with Kubernetes in production environmentsExperience deploying or operating large-scale model serving infrastructureGood understanding of distributed systems, networking, concurrency, and cloud-native architecturesExperience working with GPU computing or performance-critical applicationsPassion for solving complex infrastructure and performance engineering challengesNice to HaveExperience with NVIDIA CUDA, Triton Inference Server, or vLLMExperience serving Large Language Models in productionKnowledge of distributed inference frameworksExperience with Ray, KServe, or Kubernetes GPU OperatorsFamiliarity with observability platforms such as Prometheus, Grafana, or OpenTelemetryPrevious experience in AI Infrastructure, HPC, or AI Developer Tools companies