Originally posted 4 September 2026 by the employer — open 17 days.
About the role
This role is responsible for owning the benchmarking numbers for the T100 optical inference accelerator.
What you'll do
- Own performance and energy metrics across modeling fidelities and measured hardware.
- Produce numbers for workloads across roofline and limiter models, architecture performance models, RTL simulation, and measured competitor hardware.
- Bring up inference workloads from Hugging Face, PyTorch, and vendor stacks such as vLLM, SGLang, TensorRT-LLM, and Triton Inference Server.
- Measure competing GPUs and accelerators end-to-end, managing cloud/lab accounts, images, drivers, and run recipes.
- Report time to first token (TTFT), inter-token latency (ITL), tokens per second, tokens per second per watt, and energy using nvidia-smi, DCGM, or power capping.
- Document disagreements between RTL simulation, performance models, and measured competitor results.
What you'll need
- BS or MS in Computer Engineering, Electrical Engineering, Computer Science, or equivalent practical experience.
- 5+ years of experience in GPU performance engineering, accelerator benchmarking, HPC performance measurement, or ML systems measurement.
- Track record of building or operating benchmark harnesses that produced measured results on GPUs or accelerators.
- Hands-on experience with roofline analysis, limiter analysis, or analytical performance modeling.
- GPU performance analysis with NVIDIA Nsight Systems and Nsight Compute, covering HBM-bound versus compute-bound analysis, precision (FP16, BF16, FP8, INT8), and batching.
- Working knowledge of LLM inference stacks such as Hugging Face, vLLM, SGLang, or TensorRT-LLM, including prefill versus decode, continuous batching, and MoE.
- Proficiency in Python for harnesses, parsing, and plots, and comfort working in Linux.
- Cloud GPU operations on AWS, GCP, or Azure, including containers, instance types, drivers, quotas, and cost.
Nice to have
- Experience correlating a performance model or RTL/Verilator simulation against measured silicon or GPUs.
- GPU kernel work in CUDA, CUTLASS, or Triton, or familiarity with PyTorch internals.
- Familiarity with current inference-serving internals such as PagedAttention, FlashAttention, speculative decoding, and disaggregated prefill.
- Distributed inference experience covering collectives, all-reduce, NCCL, NVLink, and InfiniBand, or work with MLPerf or production inference benchmarking pipelines.
Skills: performance benchmarking, AI inference accelerator, silicon photonics, GPU performance engineering, LLM inference stacks
This role has been open 17 days — well below the 67-day median for AI/ML Hardware Engineering roles.
AI/ML Hardware Engineering · AI ML Hardware
|
Open roles in category
516
|
Median days open
67 d
|
Median salary
$232k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | Neurophos | All employers we track in this specialty (516 roles · 76 employers) |
|---|---|---|
| Open roles in this specialty | 5 | 516 |
| Open roles in the wider Software, Firmware & Systems family | 10 | 5829 · 143 employers |
| Median days open | 17 d | 67 d (−50 d vs this employer) |
| Median salary (USD postings) | — | $232k |
Skills observed across this category: performance benchmarking, AI inference accelerator, silicon photonics, GPU performance engineering, LLM inference stacks
Who's hiring in this category
- Qualcomm · 106 open roles · median 99 d
- NVIDIA · 85 open roles · median 67 d
- AMD · 43 open roles · median 62 d
- Micron Technology · 32 open roles · median 39 d
- Mobileye · 25 open roles · median 53 d
- Analog Devices · 14 open roles · median 24 d
How we counted: 516 open AI/ML Hardware Engineering (AI ML Hardware) roles from 76 employers tracked in the SemiconductorJobs index, counted 21 Sept 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.