أبلاي إيدج ابدأ البحث عن عمل

Machine Learning Engineer

ConnexAI · Manchester Area, United Kingdom

قدّم وتابع مع أبلاي إيدج
Build Low Latency Conversational AI SystemsWe are building real-time conversational AI systems built on top of large language models, speech AI, and agentic workflows. Our platform combines ASR, LLMs, and TTS into production-grade AI systems used globally across enterprise environments where latency, reliability, and scalability matter.We are hiring a Machine Learning Engineer to build low-latency production systems for our LLM team. This role is centred around writing scalable code that enables real-time conversational AI to perform reliably under heavy production workloads.You’ll work closely with our LLM and speech teams to solve challenges around inference speed, concurrency, request handling, GPU performance, distributed systems, and real-time response streaming.What you’ll doBuild and optimise low-latency LLM systems for real-time conversational AIWrite production-grade Python code focused on performance, scalability, and reliabilityDesign systems capable of handling large volumes of concurrent real-time requestsSolve engineering challenges around batching, request scheduling, queue management, streaming responses, and distributed workloadsImprove inference speed, GPU memory usage, and overall system responsivenessDeploy and optimise open-source LLMs using tooling such as vLLM, TensorRT-LLM, Triton, SGLang, CUDA, or similar technologiesBuild scalable orchestration layers and ML pipelines around LLM systems, including RAG and agentic workflowsDevelop backend inference services and APIs for production AI systemsProductionise new model capabilities and features for real-world customer use casesWhat we’re looking forStrong experience writing production-grade software for machine learning systemsStrong Python engineering skillsExperience building low-latency or highly concurrent systemsStrong problem-solving ability and enjoyment of building systems from the ground upExperience with distributed systems, parallel workloads, and performance optimisationExperience working with inference tooling such as vLLM, TensorRT, Triton, CUDA, ONNX, or similar technologiesExperience building scalable backend services or ML systems used in productionUnderstanding of real-time systems and performance-focused engineeringStrong communication skills and ability to work closely with engineers and researchersWhy this role?You’ll work on designing and building low-latency conversational AI systems capable of serving large volumes of concurrent real-time requests. The role focuses on solving difficult engineering challenges around inference speed, reliability, concurrency, GPU performance, and scalable production AI systems.