Skip to main content
Qualcomm

Senior/Staff AI Model Optimization Architect (Hsinchu/Taipei) - Hsinchu City, Taiwan

Qualcomm Hsinchu City, Taiwan; Taipei, Taipei City, Taiwan 5 days ago
AI/ML Hardware

Originally posted 30 September 2026 by the employer.

This role focuses on optimizing LLM/VLM/diffusion inference on Qualcomm accelerators.

About the role

This role involves leading end-to-end AI model transformation and optimization for LLMs, VLMs, diffusion, and multimodal models. The scope spans Day0 enablement through production deployment, with an emphasis on scaling optimizations to future architectures for Qualcomm inference accelerators.

What you'll do

  • Architect and deliver model optimization strategies that transform PyTorch models for efficient inference on Qualcomm accelerators.
  • Drive graph capture and deployment using PyTorch, ONNX, and torch.compile, including model rewrites and graph-level transformations.
  • Design and implement fusion kernels using DSL based approaches (e.g., Triton), enabling fused operations and performance critical algorithmic rewrites.
  • Partner with compiler, performance, and accuracy teams to co-design lowering strategies, kernel fusion, layout decisions, and runtime integration.
  • Profile and optimize LLM/VLM/diffusion inference for throughput and latency across batch sizes, sequence lengths, and serving modes.
  • Own transformer specific optimizations including KVcache management, decoding behavior, and long context performance.
  • Enable and optimize continuous batching (dynamic/iteration-level scheduling), understanding its impact on memory, scheduling, and tail latency.
  • Architect and scale distributed inference strategies (e.g., sharding and parallelism) across multi-core and multi-device systems.
  • Establish reusable approaches to scale model optimizations to new hardware architectures, creating robust patterns and tooling.
  • Debug complex performance or stability issues to root cause and drive production ready solutions.

What you'll need

  • Expert level expertise in PyTorch and inference focused model optimization.
  • Strong Python engineering skills.
  • Hands on experience with torch.compile / TorchDynamo or related graph capture and compilation workflows.
  • Deep understanding of transformer architectures, attention mechanisms, MoEs, and performance trade-offs.
  • Practical experience with KVcache behavior, serving time optimizations, and memory/performance tradeoffs.
  • Strong foundation in computer architecture, ML accelerators, and distributed systems.
  • Proven ability to lead cross-functional technical efforts and influence design decisions.
  • MS in Computer Science, Machine Learning, Computer Engineering, or Electrical Engineering, or equivalent experience.

Nice to have

  • Experience developing fusion kernels using Triton or similar DSLs, and collaborating with ML compiler teams.
  • Familiarity with LLM serving stacks and continuous batching systems.
  • Background in numerical methods, performance/accuracy trade-off analysis, or evaluation frameworks.
  • PhD in a relevant field.

Skills: AI Model Optimization, PyTorch, Qualcomm accelerators, LLM/VLM/diffusion inference, transformer architectures

Market context

This role has been open 2 days — well below the 73-day median for AI/ML Hardware Engineering roles.

AI/ML Hardware Engineering · AI ML Hardware

Open roles in category
510
Median days open
73 d
Median salary
$225k
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
How Qualcomm compares in AI/ML Hardware Engineering hiring
Metric Qualcomm All employers we track in this specialty (510 roles · 76 employers)
Open roles in this specialty 104 510
Open roles in the wider Software, Firmware & Systems family 691 5930 · 145 employers
Median days open 106 d 73 d (+33 d vs this employer)
Median salary (USD postings) — $225k

Skills observed across this category: AI Model Optimization, PyTorch, Qualcomm accelerators, LLM/VLM/diffusion inference, transformer architectures

Who's hiring in this category

  • Qualcomm (this employer) · 106 open roles · median 106 d
  • NVIDIA · 84 open roles · median 72 d
  • AMD · 42 open roles · median 66 d
  • Micron Technology · 31 open roles · median 56 d
  • Mobileye · 20 open roles · median 69 d
  • NXP Semiconductors · 15 open roles · median 59 d

How we counted: 510 open AI/ML Hardware Engineering (AI ML Hardware) roles from 76 employers tracked in the SemiconductorJobs index, counted 2 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.

Apply now
Hsinchu City, Taiwan; Taipei, Taipei City, Taiwan
On-site
5 days ago

Share this job