Research Engineering Manager, AI Agent Harness (Product R&D)
Stealth Startup ยท Abu Dhabi Emirate, United Arab Emirates
Apply & track with Apply Edge๐๐ฏ๐ ๐๐ต๐ฎ๐ฏ๐ถ (๐จ๐๐) ยท ๐๐ผ๐๐ป๐ฑ๐ถ๐ป๐ด ๐ง๐ฒ๐ฎ๐บ, ๐ฅ&๐ ยท ๐๐ , ๐๐ฑ ยท ๐ ๐๐๐ฌ๐ฏ.๐ญ.๐ฌHands-on Research Engineering Manager to establish the R&D team that builds an ๐๐ ๐ฎ๐ด๐ฒ๐ป๐ ๐ต๐ฎ๐ฟ๐ป๐ฒ๐๐ and evaluation regime for ๐ฐ๐ผ๐ป๐๐ฟ๐ผ๐น๐น๐ถ๐ป๐ด ๐๐ฒ๐ฏ ๐ฏ๐ฟ๐ผ๐๐๐ฒ๐ฟ๐ with open-weight LLMs. Greenfield.THE TECHNICAL CHALLENGEWe build agents on open-weight LLMs that operate web applications through the browser. The engineering work is the harness around the model: the abstraction that lets a pipeline or another agent drive a browser agent from the terminal, and the evaluation regime across levels of abstraction, from single actions ("find the element matching this description") to whole tasks, run against simulated applications before a live one. Agents maintain the evaluation suite; humans define its schema. Progress is measured against benchmarked results. You lead the team that builds this with the ability to write code.KEY RESPONSIBILITIES
- Set the technical direction of the harness and the interfaces between its levels of abstraction
- Establish the evaluation regime and the fortnightly score, and hold the team to measured results
- Run the R&D team on a two-week rhythm: planning, review with working software, retrospective, written post-mortems for failed experiments and incidents
- Coach, level and hire the engineers; interface with product management on the backlog and with the platform team on what the harness exposesDESIRED QUALIFICATIONS
- Managed both a research-cadence team and a sprint-cadence team, and can articulate the interface between them
- Built an agent harness or an evaluation harness for computer-use systems, agents or reinforcement learning, with an evaluation culture recognised outside the team, and can say which abstraction was right and which was wrong
- Has taken apart the harnesses of coding agents (OpenHands, SWE-agent, Aider, DeepSeek Harness or comparable) and can explain how they work in plain terms
- Has fine-tuned an LLM, from data preparation to evaluation on a held-out setEXPECTED QUALIFICATIONS
- T-shaped: deep in one domain, with working breadth in a neighbouring one
- Structures a large, incompletely specified problem and drives it to a working result independently
- Managed a team building on LLMs or machine-learning models in production, with evaluation pipelines as a management requirement and a clear account of how quality was gated
- Turns research or prototype code into reference implementations others reuse, using AI coding tools daily and verifying their outputโโHOW WE WORK
- Product engineering: we own what we build and run it in production
- Small teams, two-week cycles, working software at every review
- AI coding tools are part of the standard workflowWHAT WE OFFER
- Founding-team scope with a direct line to the Product CTO
- A supported track to the UAE Golden Visa (ten-year residency) for AI and technology talent, and relocation support including the UAE AI Specialist visa
- MacBook Pro M5 Max with 64 GB RAM or more, gear, Nvidia B200s and a serious budget for AI tokens, credits and tooling
- Up to six weeks per year working remotely from anywherePROCESS
- Introductory call with our recruiter
- Technical conversation with the Product CTO
- Practical session; the format is agreed with you
- In exercises, AI tools are allowed and expected (no LeetCode)WHO WE ARENew product organisation backed by a semi-government in Abu Dhabi. International, ex-FAANG team. Completely greenfield, with a modern tech stack.โโREQUIREMENTS TO BE CONSIDERED
- Clear written and spoken English; has reported technical progress to non-technical executives on a fixed cadence
- 6+ years in engineering, of which 2+ leading a team of engineers where work was accepted on measured results
- Still coding, able to review TypeScript and Python; own code operated in production at a product company or a startup
- Bachelor's degree in any field, or self-taught with a track record of open-source contributionsRELATED TECHNOLOGIES AND CONCEPTS
- Agent harnesses and frameworks: OpenHands, SWE-agent, Aider, DeepSeek Harness, Hermes Agent, LangGraph, smolagents
- Programmatic prompt and pipeline optimisation: DSPy, TextGrad, Ax
- Browser automation and computer use: Playwright, Chrome DevTools Protocol, Stagehand, Browser Use, Steel Browser
- Evaluation and observability: evaluation harnesses, LLM-as-judge, trajectory evaluation, Inspect, Langfuse, OpenTelemetry; benchmarks such as Online-Mind2Web, WebArena, OSWorld
- Models and serving: open-weight LLMs, vLLM, SGLang
- State and orchestration: state machines, event sourcing, XState, Trigger.dev, Temporal, Kubernetes, TypeScript, Python