Principal Machine Learning Engineer
Harnham · San Francisco Bay Area
Apply & track with Apply EdgeTitle: Principal Machine Learning EngineerLocation: Bay Area (5-days on-site)Compensation: up to $500k + Bonus + EquityWe’re partnering with a globally recognized enterprise investing heavily in AI, building one of the most widely adopted internal LLM platforms at scale. The platform is used by a large global workforce and supports use cases such as chatbots, agent workflows, code generation, and data interaction, helping teams move from idea to production significantly faster.This is a high-impact, highly visible role at the intersection of research and production, where you will help design and scale LLM-powered systems and reusable APIs used across the organization. The team operates in a fast-paced, execution-focused environment, taking problems from concept to production quickly with real adoption at scale.What You’ll DoDesign and deploy scalable LLM-powered systems and reusable backend APIsBridge cutting-edge AI research with production-grade engineeringBuild agent-based systems, RAG pipelines, and workflow automation toolsDevelop capabilities across chatbots, text-to-code, and data interaction layersPartner with cloud and SRE teams to deliver robust, scalable architecturesImplement distributed systems for training and inference (e.g., Ray, DeepSpeed)Drive best practices for ML systems, deployment, and performance optimizationRequirementsPhD in Computer Science, Mathematics, Statistics, or related fieldStrong experience as an IC in Machine Learning EngineeringProven track record in building and scaling ML/AI systems in productionDeep understanding of ML fundamentals (NLP and/or Computer Vision)Experience with distributed systems (Ray, Horovod, DeepSpeed, etc.)Strong software engineering and system design fundamentalsExperience owning ML services end-to-end in enterprise environmentsAbility to align technical work with business outcomes and communicate with stakeholdersNice to HaveExperience with ML pipelines (Kubeflow, DVC, Ray)Familiarity with microservices (gRPC, GraphQL)Experience with LLM optimization (fine-tuning, quantization, PEFT)Knowledge of advanced prompting strategies (Chain-of-Thought, etc.)If you're interested in working on AI systems at a massive scale, with real adoption and immediate impact, this is a rare opportunity to do so within a highly respected, well-resourced environment.