CUDA / GPU Performance Engineer (Kernel Optimization)
Gramian Consulting · Brazil
قدّم وتابع مع أبلاي إيدجAbout UsGramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.Role OverviewWe are looking for experienced CUDA and GPU performance engineers to analyze, profile, and optimize high-performance kernels and supporting C++ code. The role combines CUDA optimization, GPU profiling, C++, shader development, and performance analysis across different GPU architectures. No prior AI experience is required; strong systems and GPU engineering expertise is the key requirement.CONTRACT: Freelance contractor, paid per completed taskCOMMITMENT: Flexible, based on available tasks and project demandLOCATIONS: Fully remote - GLOBALPROCESS: Application review, technical assessment, and onboardingHOURLY RATE: $60-$100/hResponsibilitiesAnalyze and optimize CUDA kernels for throughput, latency, and hardware utilizationProfile GPU workloads to identify compute, memory, synchronization, and execution bottlenecksDevelop and implement targeted kernel optimization strategiesRefactor C++ and CUDA codebases for performance, maintainability, and portabilityEvaluate kernel behavior across different GPU architectures and hardware generationsDevelop or adapt shader and compute workflows using GLSL and WebGPUUse GPU profiling tools to validate improvements and compare performanceDocument optimization approaches, benchmarks, findings, and performance gainsContribute technical input to GPU architecture and performance-design discussionsEvaluate emerging GPU programming techniques and apply relevant improvementsRequirementsStrong professional experience with CUDA programming and GPU kernel optimizationAdvanced proficiency in C++, ideally in high-performance or systems programming environmentsProven experience profiling and tuning GPU workloads for performanceHands-on experience with GPU profiling tools such as NVIDIA Nsight or comparable toolsStrong understanding of GPU architecture, memory hierarchy, parallel execution, and synchronizationExperience analyzing performance across different GPU hardware generationsHands-on experience with GLSL and/or WebGPU for shader or compute developmentAbility to document performance findings and technical decisions clearly in English