Member of Technical Staff, Compilers
Acceler8 Talent · San Francisco Bay Area
Apply & track with Apply EdgeMember of Technical Staff — CompilersLocation: San Francisco, CAEmployment Type: Full-timeWorkplace: On-siteTeam: EngineeringAbout the CompanyWe are building the execution layer for the next generation of AI infrastructure.As AI workloads become larger, more dynamic, and more hardware-intensive, the answer is not simply deploying more GPUs. The harder problem is making increasingly diverse compute architectures work together efficiently at production scale.Our platform intelligently partitions, optimizes, schedules, and routes AI workloads across heterogeneous hardware environments. Customers interact through production-grade APIs without needing to manage hardware selection, placement, execution planning, or low-level optimization.We work with leading AI labs, hyperscalers, and AI-native companies running some of the most demanding inference workloads in the world.About the RoleEvery early hire changes the company.As an early member of the engineering team, you will help define the systems, standards, and technical culture behind a new class of AI infrastructure. This is a high-ownership role for someone who wants to work across compilers, runtimes, scheduling, memory systems, and hardware execution.Compilers sit at the center of this challenge.The performance gains unlocked at this layer compound across every model, workload, and hardware target on the platform. Your work will help transform modern AI workloads into efficient execution plans that run across diverse accelerator architectures in production.This is not a traditional compiler role.You will not be building a language compiler in isolation. You will be building the systems that determine how AI workloads are partitioned, optimized, scheduled, and executed across the next generation of AI infrastructure.You will work across MLIR transformations, execution planning, runtime optimization, scheduling, memory movement, kernel orchestration, speculative decoding optimization, and serving infrastructure for production-scale AI workloads.What You’ll DoIn your first 12–18 months, you will:Build compiler and runtime infrastructure that improves latency, throughput, and efficiency for large-scale AI inference workloadsDesign execution strategies that intelligently partition and coordinate workloads across heterogeneous hardwareDevelop compiler optimizations spanning IR transformations, scheduling, memory movement, and kernel orchestrationImprove serving performance for modern model architectures and advanced inference techniquesPartner with kernel, runtime, and distributed systems engineers to optimize end-to-end executionInfluence the architecture of an execution platform that will shape how AI workloads are deployed over the next decadeYou May Be a Fit If You HaveStrong systems and performance engineering fundamentalsExperience building compiler systems, compiler-adjacent infrastructure, runtime systems, or execution platformsExperience implementing IR transformations, compiler passes, lowering logic, code generation, or execution planning systemsComfort reasoning about execution behavior, memory systems, scheduling, hardware utilization, and performance tradeoffsStrong software engineering skills in C++ and/or PythonA bias toward profiling, measurement, and rigorous performance validationStrong Candidates May Also HaveExperience with MLIR, LLVM, XLA, TVM, Triton, or similar compiler/runtime infrastructureExperience optimizing ML inference, model serving, or production AI workloadsFamiliarity with runtime systems, kernel dispatch, launch APIs, memory allocators, or execution enginesExperience working with GPUs, AI accelerators, or heterogeneous hardware systemsExperience profiling and debugging performance-critical systemsFamiliarity with workload partitioning, scheduling, kernel-level optimization, or multi-accelerator executionExposure to speculative decoding, low-latency inference, or distributed serving systemsWhy This Role MattersMost AI infrastructure companies are focused on deploying more compute.We are focused on making every unit of compute more useful.The next decade of AI will be defined not only by better hardware, but by the software systems that determine how effectively workloads execute across that hardware. Compiler, runtime, and orchestration systems built today will shape how AI workloads are deployed across datacenters for years to come.As an early engineer, you will have significant ownership, work alongside deeply technical teammates, and help build the infrastructure layer that enables the next generation of AI systems.