Apply Edge Start your job search

Principal AI Engineer

Prospex Development · Al Khobar, Eastern, Saudi Arabia

Apply & track with Apply Edge

Al Khobar, KSA | Full-time | Principal individual contributor | Hands-on technical leadershipWhy this roleMillions of people work on large industrial and construction sites where a missed hazard can cost a life. Our Client’s safety platform is mandated on some client sites — a detection we get wrong has physical and contractual consequences, not just a bad metric. You will shape the AI systems that turn live site video into timely, trustworthy safety insight, built to hold up under real field conditions rather than benchmark conditions.About the roleThis is a principal-level individual-contributor role for an engineer who combines deep computer-vision and applied-AI expertise with broad technical influence. You will set architecture and quality standards, solve the hardest technical problems, and stay directly involved in implementation, experimentation, and production operation.You will guide senior engineers through design reviews, pairing, and technical mentorship, but you will not be responsible for hiring, performance management, or running squad ceremonies. Your authority comes from technical judgment, high-leverage contributions, and the ability to align teams around a sound engineering direction.This is a deliberately narrow and deep role: video is the product, not a feature of it. Our perception stack runs vision-language and computer-vision analysis over live CCTV, fuses spatial reasoning, and maps what it sees to OSHA-aligned hazard categories.What success looks like — first 6 monthsYou can explain the perception architecture end to end, including its highest-risk technical assumptions, its quality gaps, and the edge-versus-cloud trade-offs it currently makes.A measurable evaluation baseline and an automated regression gate are in place for detectors, trackers, multimodal models, prompts, and agent behaviour. Today our evaluation rigour lives in an offline, manually-run process — no accuracy regression test blocks a merge. Building that gate is one of the clearest mandates of this role, and you should expect to build it rather than inherit it.You have shipped at least one material improvement to model quality, latency, reliability, or operating cost, and validated it under representative site conditions — with the before-and-after numbers to show it.The team is using clear architecture decisions, reusable technical patterns, and production feedback to make faster, safer changes.Product, Field Engineering, and leadership trust you as the technical authority on the AI system's behaviour, its limitations, and its roadmap.What you will ownTechnical direction for the perception stack — including where vision-language and multimodal models add value, and where a lighter purpose-built detector is the right call on cost, latency, and annotation burden.The end-to-end model lifecycle — problem framing, data strategy, experimentation, fine-tuning, evaluation, deployment, monitoring, and continuous improvement.System-level quality across edge and cloud — balancing accuracy, latency, cost, throughput, privacy, and intermittent site connectivity.Responsible-AI and safety controls — confidence policies, human review for high-severity events, explainability, audit trails, and fail-safe behaviour. High-severity events need a fast path; that path trades latency against safety, and you will own where that line sits.Technical coherence across computer vision, agentic AI, video infrastructure, device integration, APIs, and production observability.What you will doDesign and build production perception systems for object detection, segmentation, tracking, event understanding, and open-ended visual reasoning — including spatial and depth reasoning, not just 2D boxes.Develop agentic AI systems that reason over live or recorded video, use tools and APIs, hold state across a shift, and operate safely in closed-loop workflows with cameras and sensors.Create representative evaluation sets and regression gates using task-appropriate measures — precision and recall, but also calibration, agreement, and significance testing — and interrogate label quality instead of assuming your ground truth is correct.Build a data flywheel using active learning, edge-case discovery, and production review signals, so labelling effort goes where it changes outcomes.Optimise and deploy models across edge accelerators and cloud infrastructure while preserving consistent behaviour and observability.Lead architecture reviews and write decision records for consequential model, data, inference, and platform choices.Make high-leverage code contributions, prototypes, and debugging interventions in the areas carrying the greatest technical risk.Partner with Product, Hardware, Field Engineering, and site stakeholders to turn operational problems into measurable AI capabilities and honest product commitments.What you will bringA career that has taken multiple production ML or computer-vision systems from problem framing to live operation — with at least one you owned end to end, including what happened to it after it shipped.Substantial hands-on work with video — decode and streaming, video analytics, or CV/VLM over video. This is the one thing the role cannot be taught on the job here.Deep hands-on experience with PyTorch or TensorFlow across object detection, segmentation, tracking, and video understanding — across both purpose-built detectors (the YOLO family) and transformer-based ones (RT-DETR, Co-DETR, GroundingDINO or equivalent), with a view on when each is the right call.Production experience with vision-language or multimodal models — open-weight (Qwen-VL, InternVL, LLaVA or equivalent) and/or hosted (Gemini, GPT, Claude) — covering prompting, grounding, adaptation, evaluation, and integration into real visual workflows, where you owned quality, latency, and cost. Tell us which models you actually put in front of users, and what it took to make them reliable enough to trust.Experience designing tool-using or agentic AI systems that reason, call services, and take controlled actions — built with LangChain/LangGraph, Google ADK, the Claude Agent SDK, LlamaIndex, or direct tool-calling — with guardrails and human-in-the-loop escalation rather than an unbounded loop.Strong data and model-adaptation practice: dataset curation, supervised or parameter-efficient fine-tuning (LoRA/QLoRA and similar), active learning, and versioned experiments. Be ready to discuss your hyperparameter choices and how you caught overfitting.Proven edge and cloud deployment experience, with practical optimisation for latency, throughput, reliability, and cost — quoted in real memory and latency budgets, not estimates.Deep evaluation and MLOps judgement: regression testing, experiment tracking, versioning, A/B testing, drift detection, monitoring, and incident learning.Principal-level influence across teams: setting direction, resolving ambiguity, and moving decisions without organisational authority.Exceptional written and spoken English. You will need to explain a model's failure mode to an HSE manager and an architecture decision to an engineer, sometimes in the same meeting.Nice to haveRetrieval-augmented generation, vector search, or multimodal retrieval pipelines.Real-time video, streaming-camera systems, or AI that operates physical devices through standard control protocols.Edge-AI toolchains and accelerators such as TensorRT, ONNX Runtime, Jetson, or other on-device GPUs and NPUs.Self-hosted LLM/VLM serving (vLLM, NIM, or similar) and reasoning about when to move off a hosted API.Experience in safety-critical, industrial, construction, or other field-operational environments; familiarity with OSHA or equivalent EHS frameworks.Privacy-by-design for video and personal data — anonymisation, access control, retention, audit trails.How we will assess youWe score demonstrated scope and measured outcomes, not tool lists or job titles. A CV line that says "cut p95 from 4.2s to 1.1s by batching and INT8 quantisation, at a 0.6-point mAP cost we accepted" tells us far more than "expert in real-time AI." We do not score university, employer brand, employer size, or years-in-seat. Where your CV is silent on something we need, we will ask about it in the interview rather than assume the worst.The tools named above are examples of a class, not a checklist. If you did the same work with InternVL instead of Qwen-VL, or Semantic Kernel instead of LangChain, that counts identically — we assess the capability, not the vendor. The same goes for cloud: our serving runs on GCP, and AWS experience is neither required nor a differentiator. What we genuinely cannot assess is a tool name with no outcome attached to it.Why join our ClientReal-world impact — help protect people on demanding industrial and construction sites, on projects at Vision-2030 scale.Principal-level ownership — you shape architecture and engineering direction while staying close to the code and the field.Modern AI problems with real site data — live video, edge compute, multimodal models, and agentic systems, on a small senior team where the work is yours end to end.Compensation and logisticsCompetitive salary package, performance bonus, and equity participation.Relocation and visa sponsorship for candidates joining from outside KSA.Health insurance, annual flights, and standard our client benefits. [TODO: confirm the exact benefits list with People Ops before posting — this line is carried over from the source template and is not verified in the context pack.]Based in Al Khobar, KSA. Excellent written and spoken English is required; Arabic is a plus, never a requirement.We welcome applicants of all backgrounds. If your experience does not match every item, we still encourage you to apply.