أبلاي إيدج ابدأ البحث عن عمل

Research Scientist

Ghost AI · San Francisco, CA

قدّم وتابع مع أبلاي إيدج
AI Research Scientist | Frontier AI | $5 Million Raised | SF Onsite Ghost AI is working with a $5M-backed AI company building the benchmarks that determine how the world's leading AI models perform in the real world.Their work spans finance, law, healthcare, software and more, with evaluations developed in-house and used to rigorously test frontier models.Now they're looking for a Research Scientist to help push this work further.Why this opportunity?This isn't a role where you'll simply run existing benchmarks.You'll have the opportunity to build new evaluation methodologies from scratch, work directly with cutting-edge AI models, and help shape how the next generation of AI systems are tested before they're deployed in the real world.The company was founded out of Stanford NLP research, with a team whose backgrounds include NVIDIA, Meta, Microsoft, Palantir and Hudson River Trading, and 300+ citations across their published research.They're growing quickly — the team has doubled in the last six months.What you'll doEvaluate frontier models as they're released, including models such as Gemini and DeepSeekDesign and build new AI benchmarks from the ground upConstruct datasets, work with labelers and develop rigorous evaluation rubricsDevelop and improve automated methods for evaluating generated textWork closely with engineers to implement and scale evaluation systemsCollaborate with leading AI labs and enterprise customers to understand emerging evaluation needsWrite technical reports and white papers around your researchOne example of their work: Public Benefits Bench, developed with Code for America and the Center for Civic Futures, tested 459 scenarios across all 50 states and found that the best-performing model was correct only 68.1% of the time.That's the kind of research that can have real-world consequences.What they're looking forMust have:1–4 years of experience in applied AI/MLResearch experience in AI/ML, LLMs, benchmarking or evaluationStrong Python skills and ability to write clean, maintainable codeExperience with PyTorch, TensorFlow or similar deep learning frameworksComfortable working in an engineering environment — Git, code reviews, development sprints, etc.Strong communication and collaboration skillsStrongly preferred:Experience building benchmarks or evaluation methodologiesExperience at an AI startup, research lab or as an early-stage/founding engineerNLP or language-model research experienceMaster's or PhD in CS, ML or a related fieldFamiliarity with LLM infrastructure and Python web frameworks such as Django or FlaskMost importantly: they're looking for someone who wants to build and apply research, not someone whose primary motivation is publishing papers.The teamThe founding team brings experience from Stanford, Palantir, HRT, NVIDIA, Meta, Microsoft and Snapchat, alongside hundreds of research citations.You'll work alongside a highly technical team with the resources to run large-scale AI evaluation studies and collaborate with some of the most important AI labs in the industry.What's on offer?Highly competitive salary + meaningful equitySan Francisco — relocation/transportation supportHealth & dental insuranceLunch & dinner provided + snacks, coffee & drinksUnlimited PTO401(k)Friday happy hours & regular team outingsIf you're an early-career AI researcher who wants to work on problems that actually influence how frontier models are evaluated, this is worth a conversation. 👻