Lead AI Systems Engineer - Vision & LLMs
Anthroholic · India
Apply & track with Apply EdgeOur mission is to deliver hyper-accurate, structured, and instant feedback on handwritten answer booklets - analyzing structural clarity, multi-dimensional content depth, diagram integration, and strict rubric alignment.We are scaling a proprietary, self-hosted, ultra-low-latency AI engine to process high-volume handwritten assessments daily. As our Lead AI Systems Engineer, you will own the core technical strategy and execution of our custom Computer Vision and open-source LLM pipeline built entirely on bare-metal GPU infrastructure.The RoleThis is the central technical leadership role in the company. You will architect, optimize, and scale our end-to-end evaluation engine-from raw image preprocessing and custom handwriting OCR/layout analysis models to local quantized LLM reasoning and high-throughput inference serving.If you excel at optimizing CUDA workloads, fine-tuning vision-language architectures, enforcing structured LLM outputs, and building high-concurrency bare-metal systems with zero reliance on third-party API dependencies, this role is for you.Key ResponsibilitiesComputer Vision & Custom Document Intelligence EngineDesign, fine-tune, and deploy state-of-the-art open-source handwriting recognition and Vision Transformer (ViT) architectures on specialized, domain-specific datasets.Build robust, fault-tolerant image preprocessing pipelines to handle automatic page alignment, deskewing, binarization, margin isolation, and document segmentation.Fine-tune models to accurately recognize multi-lingual scripts (including English and Devanagari/Hindi), diverse handwriting styles, inline margin notes, and evaluator annotations.Implement custom visual object-detection models to identify and evaluate visual elements like flowcharts, maps, and diagrams embedded within documents.On-Premise LLM Optimization & High-Throughput ServingDeploy, optimize, and maintain open-source reasoning LLMs on bare-metal GPU infrastructure using advanced local inference frameworks.Apply quantization techniques (e.g., AWQ, GPTQ, GGUF) to maximize throughput and minimize VRAM footprint while preserving evaluation precision.Enforce strict, schema-validated structured outputs (via Pydantic/Grammar-guided decoding) to ensure reliable downstream database ingestion and rendering.Core Architecture & High-Concurrency SystemsArchitect a distributed, asynchronous job-processing backend capable of handling heavy concurrent document intake with sub-minute execution targets.Implement multi-stage confidence scoring logic and an automated routing workflow for Human-in-the-Loop quality verification.Benchmark, profile, and optimize end-to-end latency and resource utilization across the entire GPU/CPU cluster.Technical & Professional RequirementsExperience: 3+ years of production experience as an AI/ML Engineer, CV Specialist, or MLOps Lead shipping real-world systems.Computer Vision: Deep expertise in PyTorch, OpenCV, layout analysis, and vision transformer/OCR architectures.LLMs & Inference: Hands-on experience serving open-source LLMs locally using production inference engines (e.g., vLLM, SGLang, TGI). Solid grasp of PagedAttention, continuous batching, and model quantization strategies.Backend & Systems: Proficiency in Python, FastAPI, distributed task queues (e.g., Celery/Redis), relational databases, Docker, and Linux/CUDA environment optimization.Performance Mindset: Demonstrated experience tracking and optimizing core production metrics: Character Error Rate (CER), Word Error Rate (WER), GPU memory utilization, and inference throughput (tokens/sec).Preferred/Nice to have SkillsDirect experience fine-tuning OCR models for Devanagari (Hindi) or other non-Latin scripts.Prior experience managing bare-metal GPU clusters and infrastructure outside standard cloud vendor ecosystems.Background in document intelligence, automated grading engines, or high-stakes assessment platforms.What we Offer?Full Engineering Ownership: Direct control over the architecture, tech stack, and roadmap powering the company's core technology.Compute Resources: Direct access to bare-metal compute and GPU resources required to build and benchmark state-of-the-art models.Impact & Equity: High-visibility role with substantial equity ownership as we scale nationwide.