Apply Edge Start your job search

Forward Deployed Engineer - LLMOps

Systems Limited · Saudi Arabia

Apply & track with Apply Edge
ABOUTOwns production operations for LLM and agentic workloads — serving, cost, and observability for a fundamentally less predictable class of system than classical ML.KEY RESPONSIBILITIESOwn production serving and scaling for LLM/agentic workloads (inference infra, load balancing, caching)Monitor and control inference cost — token usage, retry/loop cost, model routing decisionsBuild observability for LLM-specific failure modes: hallucination rate, latency spikes, prompt driftManage model/version rollout strategy (canary releases, fallback models, A/B testing)Own incident response for LLM/agent production issuesPartner with GenAI Engineers and Agentic AI Architects on production-readiness reviewsExplain token-cost dynamics to client finance/business stakeholdersCollaborate closely with GenAI Engineers without needing a hard line between build and runSupport the practice in setting cost governance policy for LLM workloadsREQUIREMENTS & SKILLS4–6 yrs platform/MLOps engineering with hands-on LLM/GenAI production experienceDeep understanding of LLM inference economics — token costs, batching, caching, model routingExperience with LLM observability tooling (tracing, eval pipelines, prompt/version management)Familiarity with multiple model hosting platforms and their cost/performance tradeoffs — Microsoft Azure AI Foundry, AWS Bedrock, and Google Vertex AI, plus self-hosted open-source options (vLLM, TGI) as a good-to-haveExperience building canary/rollback strategies for probabilistic systemsComfortable with the higher unpredictability of agentic workloads vs. classical ML servingCost-conscious communicator — can explain a token-cost blowup to a client's finance stakeholderCollaborates closely with GenAI Engineers without needing a hard line between “build” and “run”Calm under pressure during live incidents affecting client-facing systems