Inference Engineer
techire ai · San Francisco Bay Area
Apply & track with Apply EdgeFounding Inference EngineerMost infrastructure engineers joining an AI cloud inherit a platform.Here, you’ll build it.You’ll join an early-stage AI infrastructure company building a new kind of inference cloud, turning a heterogeneous fleet of AI compute into reliable, scalable infrastructure that customers can actually consume.There’s no mature platform, large infrastructure team or established playbook waiting for you.You’ll be one of the first engineers making the architectural decisions that shape how the cloud works from day one.This is a genuinely hands-on 0→1 role. You won’t be managing a team or maintaining infrastructure somebody else designed. You’ll be building the actual product.Your focus will include:
- Designing the control plane across heterogeneous compute
- Building scheduling, routing and workload/model placement systems
- Creating customer-facing APIs and serving infrastructure
- Designing reliability, observability and failure handling from the start
- Scaling the platform as workloads and customer demand growThe interesting part is how much is still unsolved.You’ll have the freedom and responsibility to decide how these systems should work, rather than fitting into an existing architecture. Decisions you make now around scheduling, serving and reliability could remain part of the platform for years.You’ll suit this if you’ve built serious distributed systems, cloud infrastructure, control planes, schedulers or large-scale serving systems and understand what reliable production infrastructure actually requires.More importantly, you’ll have experience creating systems from scratch rather than simply operating mature ones.LLM serving experience with tools such as vLLM, TGI or Ray Serve would be useful. So would experience with GPU infrastructure or alternative AI accelerators, but none of those are essential.The sweet spot is someone who sees an unsolved infrastructure problem and thinks, “I’ll build it.”Location: San FranciscoCompensation: Top-of-market + meaningful founding-stage equityIf you’ve built distributed infrastructure at scale but want considerably more technical ownership than you’d get inside an established AI or cloud company, this is worth a conversation.All applicants will receive a response.