أبلاي إيدج ابدأ البحث عن عمل

MTS - Kernel Engineer

Acceler8 Talent · Mountain View, CA

قدّم وتابع مع أبلاي إيدج
Kernel Engineer — Compute and AcceleratorsMountain View, CAAbout the RoleAn early-stage AI hardware company is seeking a Kernel Engineer to develop and optimize specialized compute kernels for a custom accelerator platform.This role sits at the critical boundary between machine learning workloads and silicon. Your work will directly influence how efficiently the hardware executes tensor operations, moves data, and uses the underlying memory hierarchy.You will work closely with architecture, compiler, simulation, and systems teams to define the kernel programming model, implement core operations, and build the profiling workflows used to evaluate hardware and software performance.What You’ll DoDevelop and optimize compute kernels for a custom AI accelerator.Implement tensor operations, data-movement patterns, and memory-hierarchy optimizations.Build and maintain profiling infrastructure to measure kernel performance against architectural targets.Define execution and data-shuffling patterns across general-purpose control cores, tensor-processing units, and specialized compute engines.Help shape the kernel programming model, including thread execution, register-passing conventions, synchronization, and memory-management strategies.Enable end-to-end kernel execution in simulation and pre-silicon environments.Collaborate with compiler engineers on intermediate representations, lowering strategies, and kernel integration.Use kernels as validation targets for compiler and architecture development.Create technical documentation, examples, and kernel-development guides for the broader engineering team.Investigate performance bottlenecks and recommend hardware or software improvements.What We’re Looking ForStrong C and C++ programming skills with experience writing production-quality systems or performance-critical code.Deep experience with CUDA or a comparable accelerator-programming model.Strong understanding of:Warp, wavefront, or thread-group executionMemory coalescingShared or local memoryRegisters and cachesSynchronizationData localityMemory bandwidth and latencyAbility to reason about computer architecture, including pipelines, execution units, memory hierarchies, and data-movement costs.Strong performance-profiling and optimization experience.Experience identifying bottlenecks, measuring throughput and latency, and iterating until performance targets are met.Practical understanding of tensor and numerical operations, including:GEMMConvolutionAttentionReductionsScatter and gatherElementwise operationsPython experience for scripting, tooling, automation, and integration work.Ability to collaborate effectively across architecture, compiler, and hardware teams.Preferred ExperienceTriton, CUTLASS, or similar kernel-development frameworks.MLIR, LLVM, or compiler infrastructure.RISC-V, x86, ARM64, or another instruction-set architecture.High-performance computing or scientific computing.Custom ASIC, GPU, NPU, or accelerator software.FPGA development or experience reading RTL.Verilog or SystemVerilog.Architectural simulators or instruction-set simulators.Kernel DSL design.Hardware-software co-design.Relevant KeywordsKernel Engineer, Compute Kernels, Accelerator Kernels, GPU Kernels, CUDA, CUDA C++, C++, Python, Triton, CUTLASS, Tensor Operations, GEMM, Matrix Multiplication, Convolution, Attention, Reductions, Scatter/Gather, Elementwise Operations, Parallel Programming, SIMT, SIMD, Warp Execution, Wavefront Execution, Thread Blocks, Memory Coalescing, Shared Memory, Registers, Cache Optimization, Memory Hierarchy, Data Locality, Data Movement, Synchronization, Performance Profiling, Performance Optimization, Throughput, Latency, Roofline Analysis, Nsight Compute, Nsight Systems, Computer Architecture, Custom ASIC, AI Accelerator, NPU, GPU, MLIR, LLVM, Kernel DSL, Compiler Integration, Architectural Simulation, Instruction Set Simulator, RISC-V, ARM64, x86, HPC, Scientific Computing, Verilog, SystemVerilog, FPGA, Hardware-Software Co-Design, Member of Technical Staff, MTS, PMTS, Principal Member of Technical Staff