Member of Technical Staff - Hardware Bringup
Infinity Artificial Intelligence Institute · San Francisco Bay Area
Apply & track with Apply EdgeMember of Technical Staff - Hardware BringupSan Francisco, CA · On-site · Full-timeAbout InfinityInfinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.About The RoleBringing a new AI accelerator from bare firmware to running state-of-the-art open-source LLMs takes months to years today. It spans firmware, kernel drivers, toolchains, compilers, kernels, runtime, and the inference stack — almost all of it written by hand, per chip.We're building Ignition: an agentic system that compresses that to under a day, on any accelerator architecture. A coding-agent controller autonomously bootstraps every layer of the stack, validates each layer against the one below it through a strict test ladder, and exposes a standards-compatible inference interface (vLLM plugin / OpenAI-compatible API) once every gate passes. Ignition took the d-Matrix Corsair — a chip with a proprietary ISA and no existing inference ecosystem — from first hardware access to tensor-parallel matmuls across all 32 compute units in 10 hours, and to three frontier models running end-to-end in 10 days. That work is now live as the Infinity d-Matrix Cloud.As one of the first engineers building this, you'll own one or more layers of the stack and the agent that generates them. Your work will be foundational to how every new chip in the industry comes online.ResponsibilitiesDepending on your strengths, you'll own one or more of the following:Build the agent controller — the coding-agent loop that consumes hardware specs, fuzzer results, and the current test gate; writes code; runs it on the target; reads results; and iterates — plus the monitoring, escalation, and dependency logic that keeps dozens of components progressing in parallelDevelop hardware characterization probes and structured ISA fuzzing that build a behavioral model of a chip with no complete spec, populating an execution-model-neutral hardware schemaBootstrap toolchains by wrapping existing tools, generating LLVM backends, or hand-writing raw instruction encoders and ELF/flat-binary loaders when nothing else existsImplement compiler and codegen strategies that branch on the chip's actual execution model (warp-based, scalar tile mesh, dataflow, flat SIMD, analog MAC)Generate and validate the kernel library — matmul, attention (MHA/GQA/MLA, flash), normalization, RoPE, MoE, and collectives — against reference implementations and measured-peak performance gatesHandle parallelism and interconnect: topology discovery, collective-algorithm selection by latency regime, and clean degradation to a single deviceBuild the runtime and serving interface: model loading, paged-attention KV cache, continuous batching, and the vLLM / OpenAI-compatible layerMaintain the test harness — the generic, per-layer test suite every gate is written against, running across multiple hardware platforms and simulationYou May Be a Good Fit If YouHave real low-level systems experience across at least two of: firmware/bare-metal, Linux kernel/driver development (ioctl, mmap, DMA, interrupts, IOMMU), compilers/codegen (LLVM, MLIR, TableGen), GPU/accelerator kernels (CUDA, ROCm/HIP, Triton, Metal), or ML inference internals (vLLM, attention kernels, KV cache, quantization)Are comfortable working from incomplete or wrong information — reverse-engineering undocumented behavior, fuzzing an ISA, reading a datasheet that doesn't match the silicon, and building a validated model anywayHave a test-first instinct, and find it satisfying rather than tedious that every layer must be provably correct against a reference before the layer above is attemptedHave hands-on experience building with coding agents / LLMs — prompting, tool-use loops, evaluating and constraining model output, and designing systems where the model writes the code and tests catch its mistakesAre fluent in Python and at least one systems language (Rust, C, or C++)Are drawn to the hardest part of the problem and comfortable being the person who figures out what's actually happening at the lowest levelStrong Candidates May Also Have Experience WithBringing up a new accelerator, board, or ISA before — vendor-side or from the outsideContributing to LLVM, MLIR, vLLM, TVM, or a hardware vendor's compiler/runtime stackNon-GPU execution models (dataflow, wafer-scale, in-memory/analog compute, RISC-V mesh)RTL simulation (Verilator, Icarus Verilog) for validating against a chip model before hardware existsDistributed training/inference (NCCL, Megatron-LM, DeepSpeed) and collective-communication internalsLogisticsDeadline to apply: None. Applications are reviewed on a rolling basis.Location: San Francisco, CA. This role is on-site.Compensation: $200k - $420k