Originally posted 20 August 2026 by the employer.
Design and build decision-grade evaluation environments for deep learning models, including LLMs, RAG, agents, and vision models.
About the role
This role involves pioneering new methodologies for accurately assessing the performance and capabilities of deep learning models, including LLMs, RAG, agents, and vision models. You will work with powerful, enterprise-grade GPU clusters capable of hundreds of PetaFLOPS and gain early access to unreleased hardware.
What you'll do
- Design and build decision-grade evaluation environments for NVIDIA's frontier models spanning reasoning, multimodal, long-context, and agentic systems, producing auditable accuracy signals.
- Research and develop novel evaluation methodologies for emerging model families and capability domains (low-precision numerics, multi-turn agentic tasks, code generation).
- Build and operate evaluation infrastructure and pipelines including benchmark environments, regression CI systems, and statistical analysis tooling.
- Partner with model research, training, and customer teams to translate evaluation signals into concrete decisions.
What you'll need
- BS, MS, or PhD in Computer Science, Machine Learning, Statistics, or a related field.
- 6+ years of hands-on experience with LLMs, designing and running evaluations for large language models or multimodal AI systems, including experience with agentic, multi-turn, or reasoning-heavy settings.
- Strong statistical foundations: experimental design, significance testing, regression analysis, and the ability to distinguish signal from noise in benchmark results.
- Proven experience building evaluation infrastructure (pipelines, benchmark harnesses, reproducible CI systems).
- Clear, precise communicator who can translate quantitative evaluation results into decisions.
Skills: Deep Learning Engineer, LLMs, multimodal AI systems, evaluation methodologies, GPU clusters, AI revolution
This role has been open 46 days — well below the 76-day median for AI/ML Hardware Engineering roles.
AI/ML Hardware Engineering · AI ML Hardware
|
Open roles in category
504
|
Median days open
76 d
|
Median salary
$230k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | NVIDIA | All employers we track in this specialty (504 roles · 75 employers) |
|---|---|---|
| Open roles in this specialty | 85 | 504 |
| Open roles in the wider Software, Firmware & Systems family | 971 | 5996 · 146 employers |
| Median days open | 71 d | 76 d (−5 d vs this employer) |
| Median salary (USD postings) | — | $230k |
Skills observed across this category: Deep Learning Engineer, LLMs, multimodal AI systems, evaluation methodologies, GPU clusters, AI revolution
Who's hiring in this category
- Qualcomm · 99 open roles · median 109 d
- NVIDIA (this employer) · 84 open roles · median 72 d
- AMD · 42 open roles · median 66 d
- Micron Technology · 30 open roles · median 58 d
- Mobileye · 18 open roles · median 73 d
- NXP Semiconductors · 15 open roles · median 63 d
How we counted: 504 open AI/ML Hardware Engineering (AI ML Hardware) roles from 75 employers tracked in the SemiconductorJobs index, counted 6 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.