Lead Applied Scientist
Greybridge Search & Selection · New York, United States
قدّم وتابع مع أبلاي إيدجLead Applied Scientist - Document Understanding | $400 - $500k total compWe're looking for a Lead Applied Scientist to help build the next generation of document understanding technology at a global information and technology business.The business supports legal, tax and accounting professionals with trusted information, software and AI-powered products. Its platforms rely on the ability to turn huge volumes of complex professional content into information that AI systems can retrieve, reason over and act on.That's where you come in.You'll work on some of the hardest problems in document understanding — from semantic chunking and document enrichment through to information extraction and knowledge graph construction. Your work will become foundational technology used across multiple product teams and large-scale AI applications.This is not a research-only position. You'll take ideas from experimentation through to production, working closely with engineering, product and subject-matter experts to build systems that are reliable, measurable and used by customers.What You'll Be DoingDesign and deploy semantic chunking models for lengthy, complex documents, going beyond fixed-size or paragraph-based approaches.Build document enrichment systems that classify content against legal, tax, accounting and customer-defined taxonomies.Develop LLM-based information extraction and knowledge graph pipelines, identifying and linking entities, citations, concepts and relationships across documents.Work on hierarchical and multi-label classification for complex document structures.Develop systems for entity recognition, entity linking, relation extraction and citation parsing.Apply knowledge distillation and model compression to develop smaller, production-ready models where latency and cost matter.Design synthetic data generation and annotation workflows to support model development and evaluation.Build component-level and end-to-end evaluation frameworks, combining expert annotation, synthetic data and automated evaluation.Make technical decisions around modelling, architecture, chunking, retrieval and knowledge extraction approaches.Partner with engineering to take models into production and ensure they perform reliably at scale.Mentor other applied scientists and ML practitioners and help raise the technical bar across the team.Provide technical input into the wider AI strategy and roadmap.What We're Looking ForYou'll have a strong background in applied NLP, document understanding or information extraction, with experience taking machine learning systems from research and experimentation into production.We're particularly interested in people who have hands-on experience with:Document intelligence and complex document structuresSemantic or structure-aware chunkingHierarchical and multi-label classificationInformation extraction from unstructured contentNER, entity linking and relation extractionCitation parsing and knowledge graph constructionLLM-based extraction and post-trainingRAG, retrieval or question-answering over large document collectionsKnowledge distillation, model compression or SLM deploymentSynthetic data generation and annotation workflowsDesigning evaluation frameworks for NLP or document understanding systemsYou'll also be comfortable working in Python, with production experience using technologies such as PyTorch and Hugging Face Transformers.A PhD or equivalent research background in Computer Science, AI, NLP or a related field is preferred, alongside a strong track record of applied research and production delivery.This is an opportunity to join that transformation at the technical foundation, helping build the systems that allow AI to understand and reason over some of the world's most complex professional content.