Apply Edge Start your job search

Head of GPU Cloud

Blue Signal Search · United States

Apply & track with Apply Edge

Head of GPU CloudRemote - NationwideOur client is building the software foundation for a next generation accelerated computing platform designed to make large-scale AI infrastructure easier to consume, operate, and scale. They are seeking a Head of GPU Cloud to lead the engineering organization responsible for transforming significant GPU capacity into high-performance cloud services for AI workloads. This is an opportunity to shape platform architecture, production inference, developer experiences, and engineering strategy at a stage where technical decisions will directly influence customer adoption, infrastructure economics, and long-term growth.This Role OffersOpportunity to define the architecture and operating model behind large scale AI inference services.Direct influence over GPU utilization, platform economics, customer experience, and technical strategy.Close collaboration with leaders across infrastructure, networking, product, operations, and commercial functions.A highly technical environment where software engineering intersects with accelerated computing, distributed systems, and AI infrastructure.FocusBuild and scale the engineering organization behind a nationwide GPU cloud platform, with ownership across inference services, orchestration, APIs, platform reliability, and technical execution.Set the architecture for moving AI workloads efficiently from customer request to accelerator, including routing, placement, model lifecycle management, caching, and memory aware scheduling.Lead production model serving and optimization across technologies such as vLLM, TensorRT LLM, and TGI, improving throughput, latency, availability, accelerator utilization, and cost.Own Kubernetes based GPU infrastructure strategy across workload scheduling, elasticity, observability, deployment automation, multi-tenant isolation, and operational resilience.Develop platform capabilities for hosted models, private inference environments, customized deployments, usage measurement, customer controls, and developer facing services.Design orchestration policies that match workloads to accelerators based on memory requirements, performance objectives, capacity, cluster topology, and infrastructure economics.Partner with networking, systems, data center, product, and commercial leaders to align software decisions with accelerator architecture, fabric performance, storage, and customer requirements.Establish engineering standards, service objectives, capacity planning, incident readiness, and team accountability while recruiting and developing senior technical talent.Skill Set12 or more years of progressive software engineering experience, including substantial leadership responsibility across cloud platforms, distributed infrastructure, HPC, or similarly complex production systems.5 or more years leading engineering teams responsible for business critical infrastructure, platform services, or other mission critical technology products.Demonstrated production experience operating large scale model inference using vLLM, TensorRT LLM, TGI, or equivalent serving stacks.Strong expertise in model serving optimization, including dynamic batching, decoding acceleration, reduced precision execution, compilation, memory reuse, and request scheduling.Advanced knowledge of Kubernetes and containerized infrastructure, including scheduling, elasticity, telemetry, deployment practices, and production reliability.Practical accelerator infrastructure knowledge covering GPU memory behavior, high bandwidth networking, InfiniBand, RoCE, storage performance, and cluster topology.Experience architecting highly available distributed platforms with automated resource allocation, programmatic interfaces, multi customer support, and detailed consumption measurement.Strong technical and leadership judgment with the ability to balance performance, reliability, security, customer experience, and infrastructure economics across multidisciplinary teams.Additional Experience That Stands OutLeadership experience in GPU cloud, AI infrastructure, hosted model platforms, or accelerated computing environments.Experience delivering elastic inference services, dedicated AI capacity, model customization workflows, or managed AI products.Familiarity with open model ecosystems and the operational differences among model families, serving configurations, and hardware profiles.Experience creating consistent developer interfaces across multiple model backends.Track record improving accelerator utilization, workload density, capacity forecasting, and compute economics across multiple GPU generations.Experience scaling engineering organizations in fast moving environments where software requirements and infrastructure capacity evolve together.About Blue Signal:Blue Signal is an award-winning, executive search firm specializing in various specialties. Our recruiters have a proven track record of placing top-tier talent across industry verticals, with deep expertise in numerous professional services. Learn more at bit.ly/46Gs4yS