Originally posted 23 September 2026 by the employer.
About the role
This internship focuses on optimizing AI computing software for NVIDIA GPUs, covering areas like LLM inference, compiler graph transformations, and CUDA kernel development.
What you'll do
- Build and enhance high-performance LLM inference pipelines.
- Analyze and optimize model execution, scalability, and memory usage.
- Improve graph transformations and code generation for NVIDIA GPUs.
- Develop compiler optimization passes, refine operator fusion and memory allocation.
- Design and tune GPU compute kernels and DSL implementations for deep learning operations.
- Profile, analyze, and improve CUDA kernel performance for maximum GPU efficiency.
What you'll need
- Pursuing an M.S. or Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or related field.
- Excellent problem-solving ability and curiosity for AI systems.
- Passion for GPU computing and deep learning software performance.
- Strong Python programming and experience with PyTorch.
- Solid understanding of inference and GPU acceleration.
- Proficient in C++.
- Experience in compiler or performance optimization.
- Skilled in C/C++ and CUDA or parallel programming.
- Familiarity with LLVM, MLIR and compiler.
- Understanding of computer architecture and performance profiling/analysis/optimization.
Skills: TensorRT LLM, compiler optimization, CUDA Kernels, deep learning operations, GPU efficiency
This role has been open 0 days — well below the 49-day median for Infrastructure/Platform Software roles.
Infrastructure/Platform Software · GPU Software Stack
|
Open roles in category
221
|
Median days open
49 d
|
Median salary
$221k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | NVIDIA | All employers we track in this specialty (221 roles · 19 employers) |
|---|---|---|
| Open roles in this specialty | 116 | 221 |
| Open roles in the wider Software, Firmware & Systems family | 1001 | 5948 · 143 employers |
| Median days open | 47 d | 49 d (−2 d vs this employer) |
| Median salary (USD postings) | — | $221k |
Skills observed across this category: TensorRT LLM, compiler optimization, CUDA Kernels, deep learning operations, GPU efficiency
Who's hiring in this category
- NVIDIA (this employer) · 109 open roles · median 47 d
- AMD · 57 open roles · median 43 d
- Qualcomm · 20 open roles · median 118 d
- Intel Corporation · 9 open roles · median 41 d
- Arm Holdings · 5 open roles · median 90 d
- Bolt Graphics · 3 open roles · median 49 d
How we counted: 221 open Infrastructure/Platform Software (GPU Software Stack) roles from 19 employers tracked in the SemiconductorJobs index, counted 23 Sept 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.