Artificial Intelligence Engineer (LLM)
UMI TECH · Singapore, Singapore
قدّم وتابع مع أبلاي إيدجResponsibilitiesBuild the end-to-end data pipeline for large language model (LLM) training, covering data cleaning, deduplication, format conversion, quality filtering, data mixture design, and version management.Design and dynamically adjust the data mix ratio between SFT (Supervised Fine-Tuning) training data and general Replay data.Manage the prompt pool for distillation training, including prompt collection, deduplication, and balanced sampling across different domains.Build a comprehensive data quality assurance system, establishing a two-level quality control process combining automated checks and manual sampling. The process should cover format validation, semantic deduplication, and persona consistency checks.Work closely with Agent Algorithm Engineers to convert ReAct multi-turn reasoning trajectories and Function Calling workflows into training-ready data formats.Manage training data versions and traceability, ensuring that data snapshots and change records for each training iteration are fully documented and traceable.RequirementsBachelor’s degree or above in Computer Science, Artificial Intelligence, Data Science, or a related field.Proficient in Python and familiar with data processing frameworks such as pandas; experience handling large-scale text datasets is required.Hands-on experience building data pipelines for LLM SFT/RLHF training, with a good understanding of best practices in data deduplication, quality filtering, and data mixture design.Familiar with ChatML/Chat formats, Function Calling data formats, and multi-turn conversational data structures.Experience with data quality measurement, monitoring, and data version management.Preferred / Nice-to-HaveExperience constructing training data for AI Agents, including ReAct trajectories and tool-calling chains.Experience with OPD and/or knowledge distillation-related data processing.Familiarity with inference frameworks such as vLLM or SGLang, with the ability to perform offline inference for large-scale data generation.