Apply Edge Start your job search

Senior Computer Vision & Multimodal AI Engineer

Infolabs Global · Dubai, United Arab Emirates

Apply & track with Apply Edge

Role: Senior Computer Vision & Multimodal AI EngineerExperience Level: 4+ Years

Employment Type: Full-Time (Paid Position)Location: DubaiWe’re Hiring: Senior Computer Vision & VLM Engineer (4+ Years Experience)Are you passionate about bridging the gap between classic computer vision and frontier Vision-Language Models (VLMs)?

We are looking for an experienced Senior Computer Vision Engineer to join our team and build next-generation, real-world visual perception and multimodal AI systems.If you have spent the last 4+ years working directly with camera pipelines, real-time vision processing, deep learning models, and multimodal architectures, we want to hear from you.What You’ll DoDesign & Deploy Vision Pipelines: Architect and optimize end-to-end computer vision pipelines processing multi-camera inputs in real time.Integrate & Fine-Tune VLMs: Build, fine-tune, and deploy state-of-the-art Vision-Language Models (e.g., LLaVA, Florence, Qwen-VL) for complex visual reasoning, grounding, and multimodal understanding.Edge & Cloud Optimization: Quantize, compile, and optimize models (TensorRT, ONNX, OpenVINO) for low-latency deployment on edge hardware and cloud infrastructure.Core CV Engineering: Develop robust solutions for object detection, tracking, semantic segmentation, and camera calibration under challenging real-world conditions.Cross-Functional Collaboration: Partner with data engineers, hardware specialists, and product teams to translate state-of-the-art vision research into production-grade features.Key RequirementsDomainRequired QualificationsExperience4+ years of hands-on professional experience building and deploying production computer vision systems.Deep Learning & VLMsDemonstrated experience with PyTorch/TensorFlow, fine-tuning VLMs/multimodal models, and prompt engineering/grounding.Camera & Video StreamsDeep understanding of camera hardware, frame capture, OpenCV, RTSP streams, and real-time video stream optimization.Edge & PerformanceProficiency in C++ and Python; experience optimizing models with TensorRT, ONNX, CUDA, or similar frameworks.FundamentalsSolid background in linear algebra, geometry, camera calibration, tracking algorithms, and loss function design.Preferred / Bonus QualificationsExperience with spatial computing, 3D vision, or SLAM systems.Track record of deploying models to embedded platforms (e.g., NVIDIA Jetson series).Published research or open-source contributions in computer vision, multimodal AI, or vision-language integration.What We OfferCompetitive Salary & Equity: Highly competitive compensation package commensurate with experience.Cutting-Edge Stack: Access to state-of-the-art compute infrastructure and early-stage hardware/models.Flexible Work Environment: Remote-first or hybrid options with flexible working hours.Comprehensive Benefits: Premium health, dental, and vision insurance, 401(k) matching, and continuous learning/conference stipends.How to ApplySend your resume, GitHub profile, and a brief note detailing your most impactful computer vision or VLM project to hr@infolabsglobal.ai or apply directly via the link below.#JobOpening #Hiring #ComputerVision #VLM #MultimodalAI #MachineLearning #DeepLearning #AIJobs #RemoteJobs #CPlusPlus #PyTorch