$208,000 - $333,500 USD yearly
Originally posted 1 October 2026 by the employer.
This role focuses on building and operating resilient AI platform capabilities at enterprise scale.
About the role
This role is for an Engineering Manager to lead a team responsible for building and operating resilient AI platform capabilities at enterprise scale. The team will focus on NVIDIA’s AI Platform Runtime and related production services.
What you'll do
- Lead, develop, and grow a team of SRE, platform, and software engineers for NVIDIA’s AI Platform Runtime and production services.
- Define the team’s technical strategy, priorities, and roadmap in alignment with product, platform, and business objectives.
- Guide the design and delivery of highly available, scalable, secure, and resilient distributed systems for enterprise AI agent products.
- Drive the development of AI agents, AI skills, and intelligent automation for platform operations, incident response, troubleshooting, and remediation.
- Establish measurable reliability goals and operational practices using service-level indicators, service-level objectives, error budgets, capacity models, operational health metrics, and production readiness reviews.
- Improve engineering velocity and developer experience through self-service platforms, infrastructure-as-code, standardized delivery patterns, and automation.
What you'll need
- 10+ years of experience in Site Reliability Engineering, Platform Engineering, Software Engineering, Cloud Infrastructure, or related technical fields.
- 3+ years managing or formally leading engineering teams responsible for complex production systems.
- Technical foundation in distributed systems, Linux, networking, Kubernetes, and public cloud platforms such as AWS, Azure, or GCP.
- Experience leading teams that build production software and automation using Python, Go, TypeScript, JavaScript, or Java.
- Solid understanding of observability at scale, including OpenTelemetry, metrics, logs, distributed tracing, profiling, and operational analytics.
- Experience applying SRE practices such as service-level objectives, error budgets, capacity and resource management, incident management, disaster recovery, and blameless postmortems.
Skills: Site Reliability Engineering, AI platform, Kubernetes, public cloud, distributed systems
This role has been open 8 days — well below the 51-day median for Infrastructure/Platform Software roles.
Infrastructure/Platform Software · Infrastructure Platform
|
Open roles in category
912
|
Median days open
51 d
|
Median salary
$229k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | NVIDIA | All employers we track in this specialty (912 roles · 86 employers) |
|---|---|---|
| Open roles in this specialty | 260 | 912 |
| Open roles in the wider Software, Firmware & Systems family | 966 | 6078 · 146 employers |
| Median days open | 46 d | 51 d (−5 d vs this employer) |
| Median salary (USD postings) | — | $229k |
Skills observed across this category: Site Reliability Engineering, AI platform, Kubernetes, public cloud, distributed systems
Who's hiring in this category
- NVIDIA (this employer) · 250 open roles · median 49 d
- Qualcomm · 60 open roles · median 61 d
- AMD · 51 open roles · median 28 d
- Cerebras · 40 open roles · median 74 d
- Graphcore · 38 open roles · median 101 d
- Intel Corporation · 33 open roles · median 21 d
How we counted: 912 open Infrastructure/Platform Software (Infrastructure Platform) roles from 86 employers tracked in the SemiconductorJobs index, counted 9 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.