Originally posted 21 August 2026 by the employer.
About the role
This role leads end-to-end model transformation and optimization for LLMs, VLMs, diffusion, and multimodal models on Qualcomm inference accelerators.
What you'll do
- Architect and deliver model optimization strategies for efficient inference on Qualcomm accelerators.
- Drive graph capture and deployment using PyTorch, ONNX, and torch.compile, including model rewrites and graph-level transformations.
- Design and implement fusion kernels using DSL based approaches (e.g., Triton), enabling fused operations and performance critical algorithmic rewrites.
- Profile and optimize LLM/VLM/diffusion inference for throughput and latency across batch sizes, sequence lengths, and serving modes.
- Own transformer specific optimizations including KVcache management, decoding behavior, and long context performance.
- Architect and scale distributed inference strategies (e.g., sharding and parallelism) across multi-core and multi-device systems.
What you'll need
- Expert level expertise in PyTorch and inference focused model optimization; strong Python engineering skills.
- Hands on experience with torch.compile / TorchDynamo or related graph capture and compilation workflows.
- Deep understanding of transformer architectures, attention mechanisms, MoEs, and performance trade-offs.
- Practical experience with KVcache behavior, serving time optimizations, and memory/performance tradeoffs.
- Strong foundation in computer architecture, ML accelerators, and distributed systems.
- Proven ability to lead cross-functional technical efforts and influence design decisions.
- MS in Computer Science, Machine Learning, Computer Engineering, or Electrical Engineering, or equivalent experience.
Nice to have
- Experience developing fusion kernels using Triton or similar DSLs, and collaborating with ML compiler teams.
- Familiarity with LLM serving stacks and continuous batching systems.
- Background in numerical methods, performance/accuracy trade-off analysis, or evaluation frameworks.
Skills: AI Model Optimization Architect, PyTorch models, Qualcomm accelerators, ONNX, torch.compile, fusion kernels, LLM/VLM/diffusion inference, transformer specific optimizations, distributed inference, hardware architectures
This role has been open 45 days — well below the 76-day median for AI/ML Hardware Engineering roles.
AI/ML Hardware Engineering · AI ML Hardware
|
Open roles in category
504
|
Median days open
76 d
|
Median salary
$230k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | Qualcomm | All employers we track in this specialty (504 roles · 75 employers) |
|---|---|---|
| Open roles in this specialty | 99 | 504 |
| Open roles in the wider Software, Firmware & Systems family | 695 | 5996 · 146 employers |
| Median days open | 109 d | 76 d (+33 d vs this employer) |
| Median salary (USD postings) | — | $230k |
Skills observed across this category: AI Model Optimization Architect, PyTorch models, Qualcomm accelerators, ONNX, torch.compile, fusion kernels
Who's hiring in this category
- Qualcomm (this employer) · 99 open roles · median 109 d
- NVIDIA · 84 open roles · median 72 d
- AMD · 42 open roles · median 66 d
- Micron Technology · 30 open roles · median 58 d
- Mobileye · 18 open roles · median 73 d
- NXP Semiconductors · 15 open roles · median 63 d
How we counted: 504 open AI/ML Hardware Engineering (AI ML Hardware) roles from 75 employers tracked in the SemiconductorJobs index, counted 6 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.