Apply Edge Start your job search

Senior Inference Engineer

Jobgether · Saudi Arabia

Apply & track with Apply Edge
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Inference Engineer based in Saudi Arabia.This is an opportunity to become the first dedicated engineer responsible for building and owning an inference platform from the ground up.You’ll work closely with the CTO to turn large language models into reliable, production-grade query-to-response systems running at scale.The role combines hands-on infrastructure engineering with performance optimization across GPU-based environments.You’ll work with modern serving technologies such as vLLM, SGLang, and TensorRT-LLM while solving real-world challenges around latency, cost, and throughput.As the platform matures, you’ll take increasing ownership of its technical direction and collaborate with Product on the future inference roadmap.You’ll join a remote-first, open-source-oriented startup where engineering ownership, speed, and measurable customer impact are highly valued.This role is ideal for a senior engineer who wants significant autonomy and the opportunity to define how production inference infrastructure is built and scaled.AccountabilitiesBuild and deploy production-grade LLM inference systems across one or multiple GPU machines, owning the complete pipeline from customer query to served response.Design, implement, and operate model-serving infrastructure using technologies such as vLLM, SGLang, and TensorRT-LLM.Optimize inference workloads for scale, balancing latency, throughput, reliability, and infrastructure costs.Apply techniques such as quantization, batching, caching, and intelligent request routing to improve inference performance.Develop robust production infrastructure using Python or Golang, with an emphasis on maintainable, scalable engineering rather than configuration-only work.Establish the initial inference platform in close collaboration with the CTO and take ownership of its evolution as the organization scales.Define and execute the technical roadmap for inference infrastructure, identifying opportunities to improve performance, reliability, and developer or customer experience.Partner with Product to translate evolving customer and market requirements into practical inference-platform capabilities.Evaluate emerging inference technologies and approaches and determine where they can create meaningful improvements.Collaborate with engineering stakeholders to establish reliable operational practices for production GPU infrastructure.RequirementsSignificant professional experience building and operating production software or infrastructure systems, with strong hands-on engineering capabilities.Demonstrated experience deploying and serving large language models in production, ideally using vLLM, SGLang, TensorRT-LLM, or comparable inference frameworks.Practical expertise optimizing inference workloads through quantization, batching, caching, routing, or similar techniques.Strong programming skills in Python or Golang, with a track record of writing and maintaining production-quality code.Strong understanding of production inference architectures, including the journey from user request through model execution to a reliable served response.Excellent problem-solving skills and the ability to independently investigate complex performance, scalability, and reliability challenges.Strong communication skills, with the ability to explain sophisticated technical concepts clearly to engineers, product stakeholders, and other audiences.Comfortable taking significant ownership, working with ambiguity, and making pragmatic technical decisions in a fast-moving environment.Familiarity with Docker and Kubernetes is a plus.Hands-on experience with generative AI technologies such as PyTorch or Transformers is advantageous.Knowledge of the GPU software stack, including CUDA, NCCL, drivers, and related libraries, is beneficial.Understanding of model architectures and fine-tuning techniques is a plus.Experience with NVIDIA Dynamo is an additional advantage.BenefitsCompetitive compensation package including equity.Health, dental, vision, and life insurance, with coverage for eligible dependents where available.Benefits adapted to the country of employment.Flexible working schedule focused on outcomes rather than fixed working hours.High degree of workplace flexibility, supporting remote work and changing personal circumstances.Remote-first environment with a globally distributed team.Significant ownership over the architecture, implementation, and long-term roadmap of the inference platform.Direct collaboration with senior technical leadership and Product teams.Opportunity to work on production-grade GPU and AI infrastructure at scale.Exposure to modern LLM serving, inference optimization, Kubernetes, cloud infrastructure, and open-source technologies.A culture built around ownership, action, continuous improvement, open-source collaboration, and technically ambitious engineering.Opportunity to help define emerging standards for AI infrastructure and production inference.How Jobgether WorksWe use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.We appreciate your interest and wish you the best! Why Apply Through Jobgether?Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.