Originally posted 5 October 2026 by the employer.
About the role
This co-op role involves contributing to research at the intersection of large language models (LLMs), distributed computing, and system optimization.
What you'll do
- Conduct research on scalable training and inference of large language models, focusing on ML systems and HPC techniques.
- Develop and optimize distributed training frameworks, model parallelism strategies, and efficient resource management for large-scale AI workloads.
- Explore hardware-aware optimizations, including algorithm-hardware co-optimization, sparsity-aware computation, quantization, and memory-efficient techniques for LLMs.
- Collaborate with researchers and engineers to publish findings in conferences.
What you'll need
- Currently pursuing a PhD in Computer Science, Electrical Engineering, or a related field with a focus on ML Systems, HPC, or AI Infrastructure.
- Strong background in machine learning, distributed systems, and parallel computing.
- Experience with deep learning frameworks (e.g., PyTorch, TensorFlow, JAX) and large-scale model training.
- Proficiency in Python and C++, with experience in performance profiling and optimization.
- Knowledge of GPUs, ASICs, distributed training paradigms (e.g., data/model pipeline parallelism, FSDP, ZeRO, DeepSpeed, Megatron-LM).
- Familiarity with HPC techniques, including MPI, Rcom/CUDA, RCCL/NCCL, and high-speed networking technologies.
- Prior research experience in scalable deep learning systems, large-scale LLM training, or AI acceleration.
- Experience with AI compiler optimizations (e.g., Triton, XLA, MLIR).
Skills: Large Language Model Engineer, ML Systems, HPC, distributed training frameworks, hardware-aware optimizations, AI acceleration
This role has been open 4 days — well below the 51-day median for Infrastructure/Platform Software roles.
Infrastructure/Platform Software · Infrastructure Platform
|
Open roles in category
912
|
Median days open
51 d
|
Median salary
$229k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | AMD | All employers we track in this specialty (912 roles · 86 employers) |
|---|---|---|
| Open roles in this specialty | 52 | 912 |
| Open roles in the wider Software, Firmware & Systems family | 406 | 6078 · 146 employers |
| Median days open | 29 d | 51 d (−22 d vs this employer) |
| Median salary (USD postings) | — | $229k |
Skills observed across this category: Large Language Model Engineer, ML Systems, HPC, distributed training frameworks, hardware-aware optimizations, AI acceleration
Who's hiring in this category
- NVIDIA · 250 open roles · median 49 d
- Qualcomm · 60 open roles · median 61 d
- AMD (this employer) · 51 open roles · median 28 d
- Cerebras · 40 open roles · median 74 d
- Graphcore · 38 open roles · median 101 d
- Intel Corporation · 33 open roles · median 21 d
How we counted: 912 open Infrastructure/Platform Software (Infrastructure Platform) roles from 86 employers tracked in the SemiconductorJobs index, counted 9 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.