Originally posted 5 October 2026 by the employer.
About the role
This role involves accelerating the adoption and optimization of AI training workloads on AMD Instinct™ GPUs.
What you'll do
- Bring up new training workloads, analyze performance bottlenecks, and develop innovative tooling.
- Develop and optimize large-scale AI training and fine-tuning workloads on AMD GPU platforms.
- Profile and analyze AI workloads, identifying bottlenecks across GPUs, CPUs, memory systems, networking, and communication infrastructure.
- Develop tools and workflows that automate training setup, debugging, performance analysis, and optimization using LLM-powered agents.
- Investigate and implement optimization strategies that improve training throughput, GPU utilization, memory efficiency, and scalability across distributed multi-GPU environments.
What you'll need
- Currently pursuing a PhD in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, Machine Learning, or a related technical discipline.
- Strong programming experience in Python and/or C++.
- Hands-on experience implementing, training, and debugging deep learning models using frameworks such as PyTorch, JAX, TensorFlow, vLLM, or SGLang.
- Experience with one or more of the following areas: Distributed training systems, data parallelism, tensor parallelism, pipeline parallelism, expert parallelism, context parallelism, GPU performance optimization, AI systems software, High-performance computing (HPC), large language model training and fine-tuning, Agentic AI or LLM-powered automation.
- Understanding of transformer-based architectures, mixture-of-experts models, and modern LLM training techniques.
- Experience profiling workloads using performance analysis tools such as PyTorch Profiler, ROCm Profiler, VTune, Nsight, or similar tools.
- Familiarity with distributed training technologies and communication libraries such as MPI, NCCL/RCCL, OpenMP, or related frameworks.
- Understanding of GPU architecture, memory systems, communication bottlenecks, and performance tuning methodologies.
Nice to have
- Experience identifying and resolving compute, memory, data-loading, or communication bottlenecks in large-scale AI workloads.
- Experience with ROCm, HIP, Triton, GPU kernel optimization, or AI systems software development.
- Publications in AI, Machine Learning, High Performance Computing, Computer Architecture, or related research areas.
Skills: AI Training Systems, AMD Instinct™ GPUs, ROCm Profiler, distributed multi-GPU environments, LLM-powered agents
This role has been open 4 days — well below the 51-day median for Infrastructure/Platform Software roles.
Infrastructure/Platform Software · Infrastructure Platform
|
Open roles in category
912
|
Median days open
51 d
|
Median salary
$229k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | AMD | All employers we track in this specialty (912 roles · 86 employers) |
|---|---|---|
| Open roles in this specialty | 52 | 912 |
| Open roles in the wider Software, Firmware & Systems family | 406 | 6078 · 146 employers |
| Median days open | 29 d | 51 d (−22 d vs this employer) |
| Median salary (USD postings) | — | $229k |
Skills observed across this category: AI Training Systems, AMD Instinct™ GPUs, ROCm Profiler, distributed multi-GPU environments, LLM-powered agents
Who's hiring in this category
- NVIDIA · 250 open roles · median 49 d
- Qualcomm · 60 open roles · median 61 d
- AMD (this employer) · 51 open roles · median 28 d
- Cerebras · 40 open roles · median 74 d
- Graphcore · 38 open roles · median 101 d
- Intel Corporation · 33 open roles · median 21 d
How we counted: 912 open Infrastructure/Platform Software (Infrastructure Platform) roles from 86 employers tracked in the SemiconductorJobs index, counted 9 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.