Senior Staff Performance Engineer - San Jose, California, United States
$189,000 - $301,000 USD yearly
Originally posted 29 September 2026 by the employer.
Join the AGI Computing Lab to solve complex system-level challenges posed by future AI/ML workloads.
About the role
This role involves designing and developing scalable platforms that effectively handle computational and memory requirements of AI/ML workloads while minimizing energy consumption and maximizing performance. The AGI Computing Lab conducts research and development in emerging technologies and trends across memory, computing, interconnect, and AI/ML.
What you'll do
- Build and operate AI environments that reflect production workloads, including agentic workflows, distributed inference, disaggregated serving architectures, and MoE deployments.
- Collect workload traces, runtime telemetry, and performance data across the software stack from AI applications.
- Characterize and compare workloads across environments and platforms, identifying compute, memory, communication, and scheduling bottlenecks.
- Communicate findings to hardware architects, systems engineers, and software researchers through reports, presentations, and architecture reviews.
- Define performance evaluation methodologies and benchmarking standards for adoption across hardware and software teams.
What you'll need
- 10+ years with a BS, 8+ years with an MS, or 5+ years with a PhD in performance engineering, AI systems, distributed systems, or high-performance computing.
- Ability to interpret workload traces, runtime telemetry, and performance data to identify bottlenecks and explain underlying causes.
- Knowledge of the LLM software stack, including serving and scheduling, attention and KV-cache management, kernel launch and memory-transfer overhead, and collective communication.
- Experience characterizing agentic workflows, long-context processing, MoE models, or disaggregated inference deployments.
- Experience profiling and optimizing AI workloads on NVIDIA GPU platforms using Nsight Systems and Nsight Compute.
- Experience analyzing multi-node AI deployments, including synchronization overhead, load imbalance, communication patterns, and scaling behavior.
- Experience with AI frameworks or serving systems such as PyTorch, vLLM, SGLang, TensorRT-LLM, DeepSpeed, Ray, or Megatron-LM.
Skills: AI/ML workloads, performance data, hardware architects, LLM software stack, NVIDIA GPU platforms
This role has been open 0 days — well below the 76-day median for AI/ML Hardware Engineering roles.
AI/ML Hardware Engineering · AI ML Hardware
|
Open roles in category
504
|
Median days open
76 d
|
Median salary
$230k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | Samsung Semiconductor | All employers we track in this specialty (504 roles · 75 employers) |
|---|---|---|
| Open roles in this specialty | 2 | 504 |
| Open roles in the wider Software, Firmware & Systems family | 14 | 5996 · 146 employers |
| Median days open | 18 d | 76 d (−58 d vs this employer) |
| Median salary (USD postings) | — | $230k |
Skills observed across this category: AI/ML workloads, performance data, hardware architects, LLM software stack, NVIDIA GPU platforms
Who's hiring in this category
- Qualcomm · 99 open roles · median 109 d
- NVIDIA · 84 open roles · median 72 d
- AMD · 42 open roles · median 66 d
- Micron Technology · 30 open roles · median 58 d
- Mobileye · 18 open roles · median 73 d
- NXP Semiconductors · 15 open roles · median 63 d
How we counted: 504 open AI/ML Hardware Engineering (AI ML Hardware) roles from 75 employers tracked in the SemiconductorJobs index, counted 6 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.