AI Hardware Systems Engineer, Annapurna Labs, Trainium Machine Learning Fleet Operations - Austin, Texas, United States
Originally posted 9 September 2026 by the employer — open 1 day.
About the role
This role involves deep diving into a fleet of ML servers deployed globally, focusing on system remediation, operational excellence, and customer experience for ML products. You will be a dedicated owner of an ML server platform, maximizing its health and customer experience.
What you'll do
- Debug emergent problems in GPU and server hardware.
- Run large scale experiments on a fleet of complex hardware.
- Develop data infrastructure and analyze trends.
- Develop automation software to scale operations.
- Utilize data to root cause hardware failures and identify live trends.
- Implement and improve system level testing across the product lifecycle.
What you'll need
- Experience debugging emergent problems in GPU hardware.
- Experience debugging emergent problems in server hardware.
- Proficiency in scripting languages such as Python.
- Proficiency in scripting languages such as Bash.
- Experience with large scale experiments on complex hardware.
- Experience developing data infrastructure and analyzing trends.
Skills: ML products, hardware failures, system level testing, ML server platform, machine learning accelerators, server products
This role has been open 0 days — well below the 63-day median for AI/ML Hardware Engineering roles.
AI/ML Hardware Engineering · AI ML Hardware
|
Open roles in category
508
|
Median days open
63 d
|
Median salary
$224k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | Amazon | All employers we track in this specialty (508 roles · 79 employers) |
|---|---|---|
| Open roles in this specialty | 2 | 508 |
| Open roles in the wider Software, Firmware & Systems family | 11 | 5650 · 142 employers |
| Median days open | 83 d | 63 d (+20 d vs this employer) |
| Median salary (USD postings) | — | $224k |
Skills observed across this category: ML products, hardware failures, system level testing, ML server platform, machine learning accelerators, server products
Who's hiring in this category
- Qualcomm · 107 open roles · median 98 d
- NVIDIA · 88 open roles · median 60 d
- AMD · 44 open roles · median 63 d
- Micron Technology · 27 open roles · median 47 d
- Mobileye · 21 open roles · median 44 d
- NXP Semiconductors · 11 open roles · median 64 d
How we counted: 508 open AI/ML Hardware Engineering (AI ML Hardware) roles from 79 employers tracked in the SemiconductorJobs index, counted 10 Sept 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.