Computer Vision Engineer
microTECH Global LTD · London Area, United Kingdom
Apply & track with Apply EdgeBacked by an international research team and abundant computing resources, the center focuses on core research directions including multimodal understanding and generation, vision-language large models, and embodied intelligence. This is a permanent, full-time position located in the tech hub of King's Cross, London.Key Responsibilities:Frontier Technical Breakthroughs-Develop ViT and multimodal large model architectures with improved reasoning and efficiency-Advance multimodal alignment, representation learning, and long-context modeling-Explore scalable training methods for large multimodal models-Optimize model architectures for generalization and performanceData Ecosystem Construction-Process large-scale multimodal data across images, videos, audio, and text-Build pipelines for data cleaning, filtering, annotation, and quality control-Construct and maintain datasets with versioning and reproducibility-Optimize data mixtures and sampling strategies for model training-Improve data quality through feedback-driven curation loopsMLLM Systems & Infrastructure-Build distributed training systems for large-scale multimodal models-Optimize GPU utilization, cluster efficiency, and resource scheduling-Develop open-source training frameworks for scalable model development-Engineer training, inference, and serving infrastructure-Improve scalability, stability, and performance of model systemsBusiness Value Delivery-Integrate multimodal capabilities into assistant and content generation scenarios-Translate research into production and user-facing applications-Collaborate with product and engineering teams to deploy and iterate modelsPerson Specification:Essential Requirements:-Academic Background: Bachelor’s degree or above in Computer Science, Mathematics, Statistics, or related technical disciplines.-Technical Skills: Proficient in Python programming with strong hands-on experience in PyTorch and deep learning frameworks.-Core Competencies: Strong algorithm development and implementation skills, solid mathematical and logical reasoning ability, and excellent cross-functional communication and collaboration skills.-Traits: Self-driven and highly motivated toward advancing artificial intelligence (AI), with strong resilience and the ability to tackle challenging technical problems.Desired:-Strong track record of publications in top-tier AI or computer vision conferences (CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR)-Hands-on experience in large-scale model pre-training or fine-tuning-High-impact open-source projects or internship experience in leading technology companies within CV, NLP, or multimodal domains