Apply Edge Start your job search

Product Lead, Inference

Inworld AI · Mountain View, CA

Apply & track with Apply Edge

Why Join InworldInworld is a research lab and inference provider focused on realtime AI for consumer-facing applications. We build first-party speech models, serve LLMs, and run the inference behind modular APIs designed for high-volume, realtime workloads.Hundreds of millions of users interact with Inworld powered apps every day and we serve over 10 trillion LLM tokens per month. Our models and infrastructure support consumer applications across companions, healthcare, fitness, education, media, and more. Our work spans model research, realtime inference, large-scale serving infrastructure, and the APIs developers use to bring these capabilities into production.We’ve raised more than $125M from Lightspeed Venture Partners, Section 32, Kleiner Perkins, Microsoft’s M12 venture fund, Founders Fund, Meta, Stanford, and others. Our technology has powered experiences from companies including NVIDIA, Microsoft Xbox, Niantic, Logitech Streamlabs, Wishroll, Little Umbrella, and Bible Chat. Inworld has also been recognized by CB Insights as one of the 100 most promising AI companies globally and named one of LinkedIn’s Top 10 Startups in the USA.Your ImpactYou will own Inworld’s inference products: the model serving and routing layer behind LLM for consumer apps at massive scale, where latency, reliability, and cost decide whether a product succeeds. This role is both strategic and hands on. You will set direction, make hard prioritization decisions, and personally ensure we ship products developers love and trust. You will collaborate with AI researchers and engineers to turn cutting-edge AI capabilities into reliable and scalable developer experiences. Your work will directly shape how the world serves and deploys real-time AI.In this role, you willDrive product vision and strategy for Inworld’s inference products, including first-party model serving and multi-model routing, to enable the future of consumer scale AI applicationsDefine the latency, reliability, and cost targets that matter for realtime workloads, and own the results against themShape how inference is packaged and priced, grounded in a strong understanding of serving costs and unit economicsDecide which models and serving capabilities to prioritize, and which workloads are best served on our own infrastructure versus routed to third-party providersGet hands on with code, both manually and with AI, to prototype, benchmark, analyze serving data, and ship production features and products in collaboration with engineeringWork closely with C-level executives and industry leaders to align our product vision with business goalsSynthesize partner and ecosystem feedback into clear product requirements and a prioritized roadmap, leading pre- and post-launch executionYou might be a good fit if you haveBA/BS degree or higher in Computer Science, Engineering, or a similar technical field6+ years of product management experience and an exceptional track record of building 0→1 developer products in high growth environmentsExperience owning an inference or ML serving product used by developers in productionA working understanding of how LLM serving works and the tradeoffs between latency, throughput, quality, and costComfort owning outcomes rather than just roadmaps, with strong product judgment in ambiguous problem spacesFounder level mindset with the ability to spot gaps, propose solutions, and move work forward without waiting for directionHands-on builder. You use AI coding tools to navigate the codebase, mock up features, run experiments, and pull your own dataA passion for learning and staying up-to-date with the latest advancements in AI and its applicationsIn-office location: Mountain View, California, United States. Candidates must be based in the SF Bay Area or willing to relocate (you will be working on-site in our South Bay office a few days a week).The US base salary range for this full-time position is $230,000 - $320,000 + bonus + equity + benefits.