Originally posted 30 September 2026 by the employer.
This role focuses on optimizing LLM/VLM/diffusion inference on Qualcomm accelerators.
About the role
This role involves leading end-to-end AI model transformation and optimization for LLMs, VLMs, diffusion, and multimodal models. The scope spans Day0 enablement through production deployment, with an emphasis on scaling optimizations to future architectures for Qualcomm inference accelerators.
What you'll do
- Architect and deliver model optimization strategies that transform PyTorch models for efficient inference on Qualcomm accelerators.
- Drive graph capture and deployment using PyTorch, ONNX, and torch.compile, including model rewrites and graph-level transformations.
- Design and implement fusion kernels using DSL based approaches (e.g., Triton), enabling fused operations and performance critical algorithmic rewrites.
- Partner with compiler, performance, and accuracy teams to co-design lowering strategies, kernel fusion, layout decisions, and runtime integration.
- Profile and optimize LLM/VLM/diffusion inference for throughput and latency across batch sizes, sequence lengths, and serving modes.
- Own transformer specific optimizations including KVcache management, decoding behavior, and long context performance.
- Enable and optimize continuous batching (dynamic/iteration-level scheduling), understanding its impact on memory, scheduling, and tail latency.
- Architect and scale distributed inference strategies (e.g., sharding and parallelism) across multi-core and multi-device systems.
- Establish reusable approaches to scale model optimizations to new hardware architectures, creating robust patterns and tooling.
- Debug complex performance or stability issues to root cause and drive production ready solutions.
What you'll need
- Expert level expertise in PyTorch and inference focused model optimization.
- Strong Python engineering skills.
- Hands on experience with torch.compile / TorchDynamo or related graph capture and compilation workflows.
- Deep understanding of transformer architectures, attention mechanisms, MoEs, and performance trade-offs.
- Practical experience with KVcache behavior, serving time optimizations, and memory/performance tradeoffs.
- Strong foundation in computer architecture, ML accelerators, and distributed systems.
- Proven ability to lead cross-functional technical efforts and influence design decisions.
- MS in Computer Science, Machine Learning, Computer Engineering, or Electrical Engineering, or equivalent experience.
Nice to have
- Experience developing fusion kernels using Triton or similar DSLs, and collaborating with ML compiler teams.
- Familiarity with LLM serving stacks and continuous batching systems.
- Background in numerical methods, performance/accuracy trade-off analysis, or evaluation frameworks.
- PhD in a relevant field.
Skills: AI Model Optimization, PyTorch, Qualcomm accelerators, LLM/VLM/diffusion inference, transformer architectures
This role has been open 2 days — well below the 73-day median for AI/ML Hardware Engineering roles.
AI/ML Hardware Engineering · AI ML Hardware
|
Open roles in category
510
|
Median days open
73 d
|
Median salary
$225k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | Qualcomm | All employers we track in this specialty (510 roles · 76 employers) |
|---|---|---|
| Open roles in this specialty | 104 | 510 |
| Open roles in the wider Software, Firmware & Systems family | 691 | 5930 · 145 employers |
| Median days open | 106 d | 73 d (+33 d vs this employer) |
| Median salary (USD postings) | — | $225k |
Skills observed across this category: AI Model Optimization, PyTorch, Qualcomm accelerators, LLM/VLM/diffusion inference, transformer architectures
Who's hiring in this category
- Qualcomm (this employer) · 106 open roles · median 106 d
- NVIDIA · 84 open roles · median 72 d
- AMD · 42 open roles · median 66 d
- Micron Technology · 31 open roles · median 56 d
- Mobileye · 20 open roles · median 69 d
- NXP Semiconductors · 15 open roles · median 59 d
How we counted: 510 open AI/ML Hardware Engineering (AI ML Hardware) roles from 76 employers tracked in the SemiconductorJobs index, counted 2 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.