Ai Distilation Expert
RB Labs · San Francisco Bay Area
قدّم وتابع مع أبلاي إيدجAbout the RoleWe're building a real-time generative AI system where latency is the core challenge. We're looking for a model distillation specialist to help us shrink large models into fast, deployment-ready versions without sacrificing quality.This is a hands-on role: you'll design and run distillation experiments, evaluate quality/latency tradeoffs, and ship the results to production GPU infrastructure.What You'll DoDesign and run knowledge distillation pipelines — teacher/student setups for large neural models (LLMs and/or generative vision models)Optimize for latency — reduce inference cost via distillation, quantization, pruning, and mixed precisionEvaluate rigorously — build empirical quality-vs-latency comparisons and A/B test variantsOwn training loops — write and debug full PyTorch training code, not just run existing scriptsShip to production — deploy distilled models on live GPU servers and measure real-world performanceMust Have2+ years hands-on PyTorch, including training/fine-tuning models from scratch — not just serving pretrained onesProven experience with model distillation — you've actually distilled a model and measured the resultsStrong grasp of model compression techniques: quantization, pruning, knowledge distillation, LoRA/parameter-efficient methodsSolid understanding of loss design for distillation (KL divergence, feature/logit matching, perceptual losses)Comfortable reading and extending dense, lightly-documented research codeStrong PlusExperience distilling autoregressive LLMs or speech modelsExperience with GANs, diffusion, or flow-matching modelsGPU inference optimization: torch.compile, CUDA memory management, multi-GPU setupsFamiliarity with streaming/real-time ML systemsWorking StyleWe operate like a research lab: empirical, fast iteration, incomplete docs, self-directed testing. You'll need to be comfortable designing your own experiments and pushing for more engineering structure as you go.