Applied AI Engineer – LLM Agents & Evaluation (Python)
Workling · San Francisco, CA
Apply & track with Apply EdgeApplied AI Engineer – LLM Agents & Evaluation (Python) Location: San Francisco Bay Area or New York City (on-site or hybrid); possibly fully remote within the US Compensation: $200,000 – $275,000 base + meaningful equity Experience: 3–8 years of professional engineering experience, with at least 1–2 years building LLM-based systems in production Employment type: Full-time About the opportunity We seek engineers whose core job is the AI itself, not a product that happens to use AI. These teams are building agents, copilots and automation for healthcare, finance, legal, security and developer tooling, and they need engineers who can take a model from "works in a demo" to "reliable in front of paying customers.". What you'll do - Design, build and ship LLM-powered agents and workflows to production: orchestration, tool use, memory, retrieval and fallback behavior - Own quality end-to-end: define evaluation datasets and metrics, build eval and benchmarking pipelines, and use them to drive prompt, model and architecture decisions - Build retrieval systems (RAG, embeddings, vector search, reranking) over messy real-world data, and the data pipelines that feed them - Run rapid experiments across models and providers (OpenAI, Anthropic, Gemini, open-weight models), measure the tradeoffs in cost, latency and accuracy, and productionize what wins - Work directly with customers and domain experts to turn ambiguous problems into shipped features, then iterate on real usage and failure cases - Fine-tune or post-train models when prompting and retrieval aren't enough, and know when that's the right call What we're looking for - 3+ years of professional software engineering experience with expert-level Python - Hands-on experience shipping LLM-based features or agents to production and keeping them reliable at scale - Strong grasp of agentic system design: tool/function calling, multi-step planning, state management, guardrails and failure handling - Experience building evaluation frameworks for non-deterministic systems: golden datasets, LLM-as-judge, regression tests, offline/online metrics - Practical retrieval and data experience: RAG pipelines, embeddings, vector databases, and processing unstructured text/documents - High ownership and product sense in an early-stage environment: you talk to users, move fast, and care about outcomes rather than models for their own sake Nice to have - PyTorch and hands-on fine-tuning / post-training (SFT, LoRA, RLHF/DPO) or reinforcement-learning environments - Inference optimization and serving (vLLM, SGLang, TensorRT-LLM), GPU/CUDA programming, or distributed training - Multimodal experience: vision, speech/voice or real-time audio - TypeScript/React for building the product surface around the AI, or Go for high-performance services - Kubernetes, AWS or GCP for deploying and scaling AI services - Research background (publications, PhD) or experience at a frontier lab - Fluency with AI coding tools (Cursor, Claude Code) as part of your daily workflow