Skip to main content
NVIDIA

Senior Production Engineer - DGX Cloud

NVIDIA 2 Locations Full-time 1 day ago
Embedded & Systems Software

Originally posted 6 October 2026 by the employer.

Build reliable, scalable, and safe AI services and endpoints for NVIDIA DGX Cloud.

About the role

This role involves building software and automation for NVIDIA DGX Cloud, which delivers AI services and endpoints for research and production workloads. The team focuses on large-scale distributed systems, including regional control plane services and GPU/CPU compute infrastructure.

What you'll do

  • Build and operate production software, automation, and tooling for control plane services, model deployments, and inference and agentic workloads.
  • Improve the reliability of inference and agentic platforms and services, including NVIDIA Cloud Functions, SGLang- and vLLM-based endpoints, and inference services built with NVIDIA Dynamo.
  • Use infrastructure as code and GitOps to deploy, configure, validate, upgrade, and recover services consistently across environments.
  • Define and instrument SLIs and SLOs for inference and control plane services, using error budgets to guide reliability improvements.
  • Participate in on-call and incident response, troubleshoot failures across routing, model runtimes, software, and infrastructure.
  • Collaborate with model, platform, storage, networking, security, and GPU infrastructure teams to design and operate services at scale.

What you'll need

  • 8+ years of experience building or operating production services and large-scale distributed systems, including hands-on automation.
  • Strong programming skills in Python, Go, or a comparable language, with experience developing tools for production operations.
  • Experience with infrastructure as code, configuration management, or GitOps, and with building automation for repeatable service deployments and changes.
  • Strong knowledge of Linux, Kubernetes, containers, cloud infrastructure, distributed systems, and networking fundamentals; ability to diagnose failures in production.
  • Understanding of Production Engineering principles, including SLIs, SLOs, error budgets, incident response, and reducing operational toil.
  • Experience instrumenting services and using metrics, logs, and traces to understand system behavior and improve reliability.
  • BS/MS in Computer Science or equivalent practical experience.

Skills: Production Engineer, large-scale distributed systems, Kubernetes, cloud infrastructure, automation

Market context

This role has been open 0 days — well below the 50-day median for Infrastructure/Platform Software roles.

Infrastructure/Platform Software · Infrastructure Platform

Open roles in category
912
Median days open
50 d
Median salary
$227k
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
How NVIDIA compares in Infrastructure/Platform Software hiring
Metric NVIDIA All employers we track in this specialty (912 roles · 86 employers)
Open roles in this specialty 261 912
Open roles in the wider Software, Firmware & Systems family 981 6043 · 146 employers
Median days open 44 d 50 d (−6 d vs this employer)
Median salary (USD postings) — $227k

Skills observed across this category: Production Engineer, large-scale distributed systems, Kubernetes, cloud infrastructure, automation

Who's hiring in this category

  • NVIDIA (this employer) · 258 open roles · median 45 d
  • Qualcomm · 58 open roles · median 59 d
  • AMD · 51 open roles · median 26 d
  • Cerebras · 40 open roles · median 72 d
  • Graphcore · 38 open roles · median 99 d
  • Renesas Electronics · 32 open roles · median 44 d

How we counted: 912 open Infrastructure/Platform Software (Infrastructure Platform) roles from 86 employers tracked in the SemiconductorJobs index, counted 7 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.

Apply now
2 Locations
On-site
Full-time
1 day ago

Share this job