Founding AI Engineer
Talotrace · Singapore, Singapore
قدّم وتابع مع أبلاي إيدجFounding AI EngineerLocation: Singapore. Hybrid.About the jobTaloTrace builds AI agents that test software on real devices. Point us at a web app or a mobile build and our agents explore it, work out what it should do, find the bugs, and file them. Web, Android and iOS, on emulators and physical devices.We are a small team of ambitious people who work at high iteration speed.We are looking for a Founding AI Engineer to own the model layer behind that. Our agents look at a screen, decide what to do, act, and judge what happened. Each of those steps is a model, and each can be smaller, faster, cheaper and more accurate than it is today.You will be the first person here whose whole job is the models, so how you set that work up is how it will run for a long time.Core ResponsibilitiesOwn the model layer end to end: data, training, evaluation, deployment, and what it costs to runTurn research results into production systems, and drop the ones that don’t hold upBuild evaluation that tells us the truth, including when the answer is that something didn’t workMake cost, quality and latency tradeoffs and back them with numbersSet the standards for model work here: what gets measured, what counts as a real improvement, and what must be true before a model shipsWhat You Could Own You will not build all of these at once. Expect to pick up two or three in your first year.Small models replacing large ones. Distillation, fine-tuning and quantization so the pipeline runs on a fraction of the computeGrounding and perception. Reliable element localisation across web, Android and iOS, including after a visual redesignThe judgment layer. Deciding whether a finding is real, how serious it is, and when to say the answer is unclearExploration policies. Models that cover an application the way real users move through it, rather than only the happy pathState prediction. Anticipating what a screen becomes after an action, which helps both navigation and catching unexpected behaviorEvaluation infrastructure. Offline benchmarks, regression gates on model changes, and metrics that measure what they claim toInference optimisation. Making the models themselves cheaper to serve quantization, batching, KV-cache work, on-device deployment. You set what each tier costs and how good it is; the platform side decides when to call whichTraining data pipelines. Collecting, cleaning, labelling and versioning the interaction data everything above runs onWhat We Are Looking For 4+ years in ML/AI engineering, or strong software engineering with substantial hands-on model work. The floor matters less than whether you have owned a model in production and can show itStrong Python and practical PyTorch or equivalent. You can read a paper’s repo and get it runningYou have trained or improved models yourself: supervised fine-tuning, LoRA/PEFT, RL, behaviour cloning or distillationYou have built evaluation harnesses and can tell when a metric is measuring the wrong thingProduction experience with LLM systems: structured output, tool use, failure modes, latency and costComfort with data pipelines. Collecting, cleaning and versioning training data is most of this jobYou check your results before you report them, and you say when something didn’t workYou build quick prototypes to answer a question and throw them away afterwardsClear writing. You can write up an experiment so someone else can act on itYou get a lot out of AI coding tools and your output holds up to reviewYou Will Stand Out If You Have Been an early engineer at a startup. You know what a first version looks likeWorked with vision-language models or multimodal trainingWorked on GUI agents, computer-use models, or robotic and embodied policiesDone imitation learning or behaviour cloning from human demonstration dataOptimised inference: FP8 or INT4 quantization, vLLM or TensorRT, KV-cache work, on-device deploymentUsed reinforcement learning, especially RLHF/GRPO or verifiable-reward setupsModelled human behaviour in any domain, whether game AI, recommender systems, user simulation or behavioural biometricsPublished, open-sourced or reproduced work in an adjacent areaProcess Intro call, a take-home challenge scoped to about 2 to 3 hours, then a live technical interview.Why TaloTrace? Our motto is leave quality to us.We think the next era of builders and changemakers should be able to reach far more people than the last one, and software is how they will do it. Our mission is to let them get from a problem statement to delivered software without stopping worrying about whether it works.Model changes land against real customer runs, so you find out whether something worked in days rather than quarters. Building the measurement that makes that trustworthy is part of the job.We have hired evidence of what you have built. We welcome applications from everyone and will accommodate whatever you need during the process, so just ask.How We Work Small team, fast pace, not much process. Where an experiment can settle a question we run it instead of debating it. Where it can’t, on direction, priorities, and what “good” means, we argue it out in the open, and we expect you to push back when you disagree. People here regularly work outside their main area.