Skip to main content
AMD

Senior GPU Kernel Optimization Engineer - Shanghai WFOE, China

AMD CN, Shanghai WFOE, China Full-time 13 days ago
Embedded & Systems Software

Originally posted 22 September 2026 by the employer — open 0 days.

About the role

This role involves designing, developing, and optimizing high-performance AI operators and GPU kernels for training and inference workloads. The focus includes improving kernel performance through profiling, memory optimization, kernel fusion, parallelization, and architecture-aware tuning.

What you'll do

  • Design, develop, and optimize high-performance AI operators and GPU kernels.
  • Analyze and improve kernel performance through profiling, memory optimization, kernel fusion, parallelization, and architecture-aware tuning.
  • Develop GPU kernels using CUDA, HIP, Triton, or similar programming frameworks.
  • Build AI-powered coding agents and automation workflows for kernel generation, optimization, benchmarking, and validation.
  • Collaborate with model and infrastructure teams to improve end-to-end AI system performance.

What you'll need

  • Bachelor's degree or above in Computer Science, Computer Engineering, Artificial Intelligence, or a related field, with 4+ years of relevant experience.
  • Strong programming skills in C++ and/or Python.
  • Hands-on experience in GPU kernel development and optimization using CUDA, HIP, Triton, or similar frameworks.
  • Solid understanding of GPU architecture and performance optimization, including memory hierarchy, occupancy, register pressure, and shared memory.
  • Experience with performance profiling and optimization tools such as Nsight, rocprof, or equivalent.
  • Experience with AI operators such as GEMM, Attention, MoE, Embedding, or other performance-critical workloads.

Nice to have

  • Deep experience in GPU kernel optimization, including kernel fusion, memory optimization, communication/computation overlap, performance modeling, or optimization of AI operators such as GEMM, Attention, MoE, and Embedding.
  • Experience building coding agents, code generation systems, or AI-assisted software engineering tools.
  • Experience with large-scale AI training or inference systems.
  • Experience with LLM post-training, including SFT, RLHF, DPO, GRPO, or related techniques.
  • Publications or open-source contributions in GPU optimization, AI systems, or coding agents.
Market context

This role has been open 0 days — well below the 56-day median for Systems Software Engineering roles.

Systems Software Engineering · Systems Software

Open roles in category
674
Median days open
56 d
Median salary
$205k
See the full market breakdown ▾Category comparison, and who else is hiring
How AMD compares in Systems Software Engineering hiring
Metric AMD All employers we track in this specialty (674 roles · 72 employers)
Open roles in this specialty 57 674
Open roles in the wider Software, Firmware & Systems family 399 5873 · 143 employers
Median days open 53 d 56 d (−3 d vs this employer)
Median salary (USD postings) — $205k

Who's hiring in this category

  • Qualcomm · 157 open roles · median 63 d
  • NVIDIA · 132 open roles · median 46 d
  • AMD (this employer) · 56 open roles · median 53 d
  • Broadcom · 26 open roles · median 54 d
  • Mobileye · 24 open roles · median 62 d
  • Arm Holdings · 20 open roles · median 89 d

How we counted: 674 open Systems Software Engineering (Systems Software) roles from 72 employers tracked in the SemiconductorJobs index, counted 22 Sept 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.

Apply now
CN, Shanghai WFOE, China
On-site
Full-time
13 days ago

Share this job