أبلاي إيدج ابدأ البحث عن عمل

Machine Learning Engineer — Inference Optimization

Jobgether · United Arab Emirates

قدّم وتابع مع أبلاي إيدج
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Machine Learning Engineer — Inference Optimization based in United Arab Emirates.This role offers the opportunity to optimize the performance of advanced machine learning systems used in real-world production environments.You will work at the intersection of research and engineering, transforming cutting-edge models into fast, reliable, and cost-efficient solutions.Your work will directly impact model scalability, user experience, and the efficiency of AI-powered products.You will dive deep into performance optimization, from model architecture and GPU execution to large-scale inference infrastructure.Working with talented research, infrastructure, and product teams, you will help push the boundaries of what AI systems can achieve.This position is ideal for an engineer who enjoys solving complex technical challenges and building high-performance ML systems from the ground up.AccountabilitiesAs a Machine Learning Engineer specializing in inference optimization, you will own the performance and scalability of machine learning models in production. You will combine deep ML expertise, systems engineering, and performance analysis to deliver faster, more efficient AI experiences.Optimize machine learning inference systems to improve latency, throughput, scalability, and operational cost.Profile and identify bottlenecks across GPU and CPU inference pipelines, including memory usage, kernels, batching strategies, and data flow.Implement advanced optimization techniques such as quantization, KV-cache optimization, speculative decoding, batching, streaming, and model simplification.Collaborate with research engineers to productionize new model architectures and translate experimental results into reliable systems.Build, improve, and maintain inference-serving infrastructure using modern frameworks, custom runtimes, or specialized serving solutions.Benchmark model performance across different hardware environments, including GPUs, CPUs, and cloud-based systems.Improve system reliability, monitoring, observability, and cost efficiency under real production workloads.Contribute to engineering practices that improve the quality, scalability, and maintainability of ML infrastructure.RequirementsThe ideal candidate is a technically strong machine learning engineer with experience optimizing production inference systems and a passion for high-performance AI engineering. You should enjoy working on complex technical problems, experimenting with new approaches, and taking ownership of critical systems.Strong professional experience in ML inference optimization, high-performance machine learning systems, or related areas.Deep understanding of machine learning fundamentals, including neural network architectures, attention mechanisms, memory optimization, and compute graphs.Hands-on experience with PyTorch or similar deep learning frameworks and deploying models into production environments.Experience with GPU performance optimization, including technologies such as CUDA, ROCm, Triton, or kernel-level tuning.Proven experience scaling inference systems for real users beyond research prototypes or benchmarks.Strong programming skills and the ability to work across machine learning and systems engineering domains.Ability to operate effectively in fast-paced environments with ownership, autonomy, and evolving priorities.Experience with inference frameworks such as TensorRT, ONNX Runtime, vLLM, or Triton is a plus.Familiarity with large language models, long-context inference, distributed systems, low-latency services, or hardware optimization is considered an advantage.Contributions to open-source ML systems or inference tooling are a plus.BenefitsCompetitive compensation package with meaningful equity participation.Opportunity to work on performance-critical AI systems with direct product impact.High level of ownership over infrastructure that shapes scalability and efficiency.Close collaboration with research, infrastructure, and product teams.Opportunity to work on advanced machine learning technologies and real-world AI applications.Engineering-focused culture that values technical excellence, experimentation, and quality.Flexible remote work environment.Opportunity to contribute to the growth of an innovative AI-focused organization.How Jobgether WorksWe use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.We appreciate your interest and wish you the best! Why Apply Through Jobgether?Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.