Skip to main content
Qualcomm

Senior/Staff AI Performance Engineer, Cloud AI (Hsinchu/Taipei) - Hsinchu City, Taiwan

Qualcomm Hsinchu City, Taiwan; Taipei, Taipei City, Taiwan 5 days ago
AI/ML Hardware

Originally posted 30 September 2026 by the employer.

Join the Cloud AI team to optimize LLM, VLM, and diffusion models.

About the role

This role focuses on developing hardware and software solutions for Inference Acceleration in Cloud AI, spanning from research and development to commercial deployment.

What you'll do

  • Convert, optimize, and deploy models for efficient inference using PyTorch and ONNX.
  • Analyze advanced algorithms like attention mechanisms and MoEs to identify optimization opportunities.
  • Optimize LLM, VLM, and diffusion models for inference performance, throughput, and latency.
  • Collaborate with internal compiler, firmware, and platform teams to drive customer solutions.
  • Analyze complex performance or stability issues to determine root causes.
  • Create engineering solutions for continuous insights into AI workload performance.
  • Design and implement high-level kernels, such as in Triton, for efficient low-level code generation.

What you'll need

  • Hands-on experience building and optimizing language models in PyTorch and ONNX, preferably in production.
  • Deep understanding of transformer architectures, attention mechanisms, and performance trade-offs.
  • Experience in workload mapping strategies, including sharding or various parallelisms.
  • Strong Python programming skills.
  • Proactive learning regarding the latest inference optimization techniques.
  • Understanding of computer architecture, ML accelerators, in-memory processing, and distributed systems.
  • MS in Computer Science, Machine Learning, Computer Engineering, or Electrical Engineering.

Nice to have

  • Background in neural network operators and mathematical operations, including linear algebra and math libraries.
  • Understanding of machine learning compilers.
  • Experience in converging accuracy and its evaluation methods.
  • Knowledge of torch.compile or torchDynamo.
  • PhD in Computer Science, Computer Engineering, or Machine Learning.

Skills: Cloud AI, Inference Acceleration, LLM, VLM, diffusion models, ML accelerators

Market context

This role has been open 2 days — well below the 73-day median for AI/ML Hardware Engineering roles.

AI/ML Hardware Engineering · AI ML Hardware

Open roles in category
510
Median days open
73 d
Median salary
$225k
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
How Qualcomm compares in AI/ML Hardware Engineering hiring
Metric Qualcomm All employers we track in this specialty (510 roles · 76 employers)
Open roles in this specialty 104 510
Open roles in the wider Software, Firmware & Systems family 691 5930 · 145 employers
Median days open 106 d 73 d (+33 d vs this employer)
Median salary (USD postings) — $225k

Skills observed across this category: Cloud AI, Inference Acceleration, LLM, VLM, diffusion models, ML accelerators

Who's hiring in this category

  • Qualcomm (this employer) · 106 open roles · median 106 d
  • NVIDIA · 84 open roles · median 72 d
  • AMD · 42 open roles · median 66 d
  • Micron Technology · 31 open roles · median 56 d
  • Mobileye · 20 open roles · median 69 d
  • NXP Semiconductors · 15 open roles · median 59 d

How we counted: 510 open AI/ML Hardware Engineering (AI ML Hardware) roles from 76 employers tracked in the SemiconductorJobs index, counted 2 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.

Apply now
Hsinchu City, Taiwan; Taipei, Taipei City, Taiwan
On-site
5 days ago

Share this job