Originally posted 30 September 2026 by the employer.
Join the Cloud AI team to optimize LLM, VLM, and diffusion models.
About the role
This role focuses on developing hardware and software solutions for Inference Acceleration in Cloud AI, spanning from research and development to commercial deployment.
What you'll do
- Convert, optimize, and deploy models for efficient inference using PyTorch and ONNX.
- Analyze advanced algorithms like attention mechanisms and MoEs to identify optimization opportunities.
- Optimize LLM, VLM, and diffusion models for inference performance, throughput, and latency.
- Collaborate with internal compiler, firmware, and platform teams to drive customer solutions.
- Analyze complex performance or stability issues to determine root causes.
- Create engineering solutions for continuous insights into AI workload performance.
- Design and implement high-level kernels, such as in Triton, for efficient low-level code generation.
What you'll need
- Hands-on experience building and optimizing language models in PyTorch and ONNX, preferably in production.
- Deep understanding of transformer architectures, attention mechanisms, and performance trade-offs.
- Experience in workload mapping strategies, including sharding or various parallelisms.
- Strong Python programming skills.
- Proactive learning regarding the latest inference optimization techniques.
- Understanding of computer architecture, ML accelerators, in-memory processing, and distributed systems.
- MS in Computer Science, Machine Learning, Computer Engineering, or Electrical Engineering.
Nice to have
- Background in neural network operators and mathematical operations, including linear algebra and math libraries.
- Understanding of machine learning compilers.
- Experience in converging accuracy and its evaluation methods.
- Knowledge of torch.compile or torchDynamo.
- PhD in Computer Science, Computer Engineering, or Machine Learning.
Skills: Cloud AI, Inference Acceleration, LLM, VLM, diffusion models, ML accelerators
This role has been open 2 days — well below the 73-day median for AI/ML Hardware Engineering roles.
AI/ML Hardware Engineering · AI ML Hardware
|
Open roles in category
510
|
Median days open
73 d
|
Median salary
$225k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | Qualcomm | All employers we track in this specialty (510 roles · 76 employers) |
|---|---|---|
| Open roles in this specialty | 104 | 510 |
| Open roles in the wider Software, Firmware & Systems family | 691 | 5930 · 145 employers |
| Median days open | 106 d | 73 d (+33 d vs this employer) |
| Median salary (USD postings) | — | $225k |
Skills observed across this category: Cloud AI, Inference Acceleration, LLM, VLM, diffusion models, ML accelerators
Who's hiring in this category
- Qualcomm (this employer) · 106 open roles · median 106 d
- NVIDIA · 84 open roles · median 72 d
- AMD · 42 open roles · median 66 d
- Micron Technology · 31 open roles · median 56 d
- Mobileye · 20 open roles · median 69 d
- NXP Semiconductors · 15 open roles · median 59 d
How we counted: 510 open AI/ML Hardware Engineering (AI ML Hardware) roles from 76 employers tracked in the SemiconductorJobs index, counted 2 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.