$224,000 - $356,500 USD yearly
Originally posted 15 September 2026 by the employer — open 4 days.
About the role
This role involves owning the data engine and end-to-end training verification for 3D spatial reasoning and perception capabilities within NVIDIA Cosmos models.
What you'll do
- Own the 3D data engine for Cosmos spatial reasoning, including sourcing, curating, filtering, and balancing large-scale real-world image and video corpora.
- Build annotation and auto-labeling pipelines for 3D-grounded supervision at scale, covering aspects like camera-relative 3D boxes and cross-view correspondence.
- Manage data quality end-to-end, including semantic deduplication, automated quality scoring, coverage analysis, and sampling strategies.
- Verify end-to-end model training, guard reproducibility, diagnose throughput/loss anomalies, and attribute capability changes.
- Build and operate the 3D and spatial evaluation suite, including public benchmarks such as CV-Bench, BLINK, RefSpatial, VSI-Bench, SPAR-Bench, and RoboSpatial, and NVIDIA's VANTAGE-Bench.
- Operate on large multi-node GPU clusters, tuning data throughput, sharding, and dataloader performance.
What you'll need
- MS or PhD in Computer Science, Electrical/Computer Engineering, Robotics, or a related field, or equivalent experience.
- 12+ years of experience building deep learning systems in Python with PyTorch or JAX on Linux.
- Expertise in 3D computer vision, multi-view geometry, structure-from-motion or SLAM, depth and camera pose estimation, point cloud processing, or 3D reconstruction.
- Hands-on experience with vision-language models, including building training data and evaluations.
- Experience building large-scale multimodal data pipelines: distributed video and image processing, deduplication, captioning and annotation, automated quality metrics, and dataset versioning.
- Experience running and validating large model training on multi-GPU, multi-node clusters, with knowledge of distributed training and sharding strategies such as data, tensor, and pipeline parallelism or FSDP.
- Rigorous evaluation methodology: designing benchmarks, building clean ablations, and reading results.
Skills: Deep Learning Engineer, Physical AI, 3D spatial reasoning, vision-language models, multi-node GPU clusters
This role has been open 3 days — well below the 65-day median for AI/ML Hardware Engineering roles.
AI/ML Hardware Engineering · AI ML Hardware
|
Open roles in category
518
|
Median days open
65 d
|
Median salary
$229k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | NVIDIA | All employers we track in this specialty (518 roles · 76 employers) |
|---|---|---|
| Open roles in this specialty | 85 | 518 |
| Open roles in the wider Software, Firmware & Systems family | 943 | 5796 · 143 employers |
| Median days open | 65 d | 65 d (same as this employer) |
| Median salary (USD postings) | — | $229k |
Skills observed across this category: Deep Learning Engineer, Physical AI, 3D spatial reasoning, vision-language models, multi-node GPU clusters
Who's hiring in this category
- Qualcomm · 106 open roles · median 97 d
- NVIDIA (this employer) · 85 open roles · median 65 d
- AMD · 43 open roles · median 60 d
- Micron Technology · 32 open roles · median 37 d
- Mobileye · 25 open roles · median 51 d
- Analog Devices · 14 open roles · median 22 d
How we counted: 518 open AI/ML Hardware Engineering (AI ML Hardware) roles from 76 employers tracked in the SemiconductorJobs index, counted 19 Sept 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.