Forward Deployed Engineer - AI Assurance
Systems Limited · Saudi Arabia
قدّم وتابع مع أبلاي إيدجABOUT:Owns quality for AI-native applications — functional testing plus the AI-specific evaluation (accuracy, drift, hallucination).KEY RESPONSIBILITIESBuild test plans and automation for AI-native application features (functional + AI-specific)Design evaluation harnesses for model/agent outputs — accuracy, consistency, hallucination rateRun regression testing across model/prompt/config changes to catch silent quality driftRed-team AI features for edge cases and adversarial inputs where relevantBuild automated eval pipelines integrated into CI/CDPartner with AI Architects to define testability requirements before build startsOwn the quality gate before any AI feature ships to productionCommunicate quality risk to delivery leadership in terms they can act onTrain delivery teams on AI-specific testing practicesOwn the evals framework for the practice — golden datasets, scoring rubrics, LLM-as-judge calibration, and versioned benchmarks per use caseDefine eval acceptance thresholds per engagement and gate releases on themBuild eval engineering tooling — dataset curation, trace capture, offline/online eval runs, and dashboards delivery teams can readInstrument production evals and drift monitoring, feeding failures back into the golden datasetsREQUIREMENTS & SKILLS5–9 yrs QA/test engineering, with 2+ yrs testing AI/ML-powered features specificallyStrong test automation skills (Python-based frameworks, CI/CD integration)Understands AI-specific failure modes — hallucination, bias, drift, non-determinism — and designs tests for themStatistically literate enough to interpret model evaluation metrics, not just pass/fail resultsFamiliarity with red-teaming methodologies for AI systemsClear, assertive communicator — willing to block a release over a quality concernDetail-oriented and methodical under delivery-timeline pressureCollaborative but independent — doesn't rubber-stamp under delivery pressureExplains quality risk in business-impact terms, not just technical jargonHands-on evals engineering — builds and maintains eval suites with frameworks such as OpenAI Evals, Ragas, DeepEval, LangSmith, Azure AI Foundry evaluationsDesigns golden datasets and rubrics, and calibrates LLM-as-judge scoring against human reviewUnderstands RAG and agent eval metrics — groundedness, retrieval precision/recall, task completion, tool-call correctness, cost/latencyExperience wiring evals and drift monitoring into CI/CD and production observability