Skip to main content
NVIDIA

Senior Product Manager, Inference Platform - CA, Santa Clara, United States

NVIDIA Remote friendly (US, CA, Santa Clara, United States) 13 days ago
Product & Program Management

$208,000 - $379,500 USD yearly

Originally posted 28 September 2026 by the employer.

Define and drive products and platform capabilities for large-scale model serving.

About the role

As a Product Manager for Inference Platform, you will define and drive the products and platform capabilities that enable large-scale model serving across a broad portfolio of models. You will work at the intersection of AI research, infrastructure engineering, and real user needs, shaping how inference is delivered reliably, efficiently, and at scale.

What you'll do

- Define product vision and strategy for inference platform capabilities, including APIs, capacity management, cost management, performance and optimization, and model serving infrastructure. - Translate user needs and infrastructure constraints into clear requirements and prioritized roadmaps. - Partner closely with engineering, research, and user groups teams to drive execution from concept through launch. - Be responsible for end-to-end product lifecycle for inference-related products and platform investments. - Develop deep understanding of the inference ecosystem, including model formats, serving frameworks, API formats, and relevant tradeoffs at scale. - Track and synthesize developments across the inference landscape: open source model releases, serving frameworks, competitive dynamics, and emerging use cases.

What you'll need

- 12+ years of experience with a track record of delivering complex technical products. - Bachelors degree or higher, or equivalent experience. - Strong written and verbal communication, with the ability to write clear, concise product documents, specs, and strategies. - Deep familiarity with AI/ML systems and inference serving frameworks (such as TensorRT-LLM, vLLM, or Triton Inference Server), tradeoffs involved, and what matters to model publishers, application developers and cloud operators. - Experience with inference APIs- design, versioning, performance, hardware efficiency, and developer experience. - Familiarity with open source as well as commercial model ecosystems and the different considerations each brings to a serving platform.

Skills: product vision, product strategy, inference platform, model serving, TensorRT-LLM, Triton Inference Server

Market context

This role has been open 11 days — well below the 44-day median for Product Management roles.

Product Management · AI ML Hardware

Open roles in category
66
Median days open
44 d
Median salary
$248k
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
How NVIDIA compares in Product Management hiring
Metric NVIDIA All employers we track in this specialty (66 roles · 21 employers)
Open roles in this specialty 20 66
Median days open 37 d 44 d (−7 d vs this employer)
Median salary (USD postings) — $248k

Skills observed across this category: product vision, product strategy, inference platform, model serving, TensorRT-LLM, Triton Inference Server

Who's hiring in this category

  • NVIDIA (this employer) · 20 open roles · median 37 d
  • Qualcomm · 10 open roles · median 58 d
  • AMD · 7 open roles · median 43 d
  • Micron Technology · 4 open roles · median 65 d
  • Analog Devices · 3 open roles · median 28 d
  • NXP Semiconductors · 3 open roles · median 10 d

How we counted: 66 open Product Management (AI ML Hardware) roles from 21 employers tracked in the SemiconductorJobs index, counted 10 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there. Figures refresh nightly.

Apply now
US, CA, Santa Clara, United States
Hybrid
$208,000 - $379,500 USD yearly
13 days ago

Share this job