Originally posted 21 September 2026 by the employer — open 0 days.
Optimize performance for AI Kernels.
About the role
This role involves optimizing the performance of key AI workloads and benchmarks for AMD GPUs, focusing on machine learning kernels.
What you'll do
- Identify and implement improvements to machine learning kernels for AMD GPUs, focusing on performance and power efficiency.
- Stay informed about software and hardware trends, particularly in GPU architecture and machine learning algorithms.
- Improve development workflows and CI infrastructure to enable faster and more reliable delivery.
- Design and develop innovative GPU and machine learning technologies.
- Debug and resolve existing issues, while researching and implementing more efficient alternatives.
- Build and maintain strong technical relationships with internal teams and external partners.
What you'll need
- Strong understanding of modern GPU architectures.
- 3+ years of GPU software development experience using HIP, CUDA, or OpenCL.
- 5+ years of system-level programming experience in C++ (C++17 or later preferred).
- Experience with GPU profiling, debugging, benchmarking, and performance analysis tools.
- Background in high-performance computing (HPC) or other performance-critical systems.
- Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience.
Nice to have
- Experience working directly with hardware ISA.
- Familiarity with modern machine learning frameworks such as PyTorch and MIOpen.
- Experience with tile-based programming models and frameworks (e.g., Triton, CUTLASS).
Skills: GPU architectures, HIP, CUDA, OpenCL, PyTorch, MIOpen
This role has been open 0 days — well below the 48-day median for Infrastructure/Platform Software roles.
Infrastructure/Platform Software · GPU Software Stack
|
Open roles in category
217
|
Median days open
48 d
|
Median salary
$221k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | AMD | All employers we track in this specialty (217 roles · 19 employers) |
|---|---|---|
| Open roles in this specialty | 57 | 217 |
| Open roles in the wider Software, Firmware & Systems family | 392 | 5829 · 143 employers |
| Median days open | 41 d | 48 d (−7 d vs this employer) |
| Median salary (USD postings) | — | $221k |
Skills observed across this category: GPU architectures, HIP, CUDA, OpenCL, PyTorch, MIOpen
Who's hiring in this category
- NVIDIA · 107 open roles · median 45 d
- AMD (this employer) · 56 open roles · median 43 d
- Qualcomm · 19 open roles · median 150 d
- Intel Corporation · 9 open roles · median 39 d
- Arm Holdings · 5 open roles · median 88 d
- Bolt Graphics · 3 open roles · median 47 d
How we counted: 217 open Infrastructure/Platform Software (GPU Software Stack) roles from 19 employers tracked in the SemiconductorJobs index, counted 21 Sept 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.