Originally posted 23 September 2026 by the employer.
About the role
This role involves building models and tools to predict how generative AI inference workloads perform on Tensordyne systems, from a single accelerator up through rack, pod, and cluster scale. The position focuses on hands-on engineering, including writing simulator code, running experiments, and debugging performance predictions against real hardware.
What you'll do
- Implement and extend simulation-based performance models for multimodal generative AI inference at rack, pod, and cluster scale, covering compute, memory, collective communication, and network fabric.
- Model how serving strategies (tensor, pipeline, and expert parallelism, prefill/decode disaggregation, batching, and KV-cache placement) interact with Tensordyne silicon and fabric topology, and measure the effect on latency, throughput, and cost per token.
- Build trace-capture and replay tooling that records real execution from our inference runtime and replays it under hypothetical silicon, system, and network configurations.
- Model collective communication on multi-hop scale-out fabrics, including implementing custom collective algorithms designed for our topology.
- Run calibration experiments on Tensordyne hardware as systems come up, compare them against model predictions, and fix the sources of error.
- Run design-space sweeps and write up clear analyses that architects and engineering teams use in ASIC, fabric, and system configuration decisions.
What you'll need
- Hands-on experience building performance models, simulators, or analytical tools for ML workloads, distributed systems, or computer architecture.
- Solid understanding of distributed ML execution, including parallelism strategies, collective communication (All-Reduce, All-Gather, All-to-All, etc.), and how they scale.
- Working knowledge of system architecture across compute, memory, interconnect, and networking, and the ability to reason about bottlenecks between them.
- Experience comparing model predictions against real measurements, and debugging where they diverge.
- Strong programming skills in C++ and Python, with clean, testable, maintainable code.
- MS or higher in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
Nice to have
- Familiarity with LLM inference serving: batching, KV-cache management, disaggregated prefill/decode, and latency/throughput trade-offs.
- Experience modeling or benchmarking collective communication libraries (NCCL, RCCL, or similar) on real clusters.
- Background in data center or HPC networking: topologies, RDMA/RoCE, and congestion behavior.
- Experience profiling ML workloads on accelerators (GPUs, TPUs, or custom ASICs).
- Exposure to hardware/software co-design or early-stage architecture evaluation.
- Publications or open-source contributions in ML systems, architecture, or networking.
Skills: generative AI inference, performance modeling, ML workloads, distributed ML execution, system architecture, C++
This role has been open 0 days — well below the 69-day median for AI/ML Hardware Engineering roles.
AI/ML Hardware Engineering · AI ML Hardware
|
Open roles in category
517
|
Median days open
69 d
|
Median salary
$232k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | Tensordyne | All employers we track in this specialty (517 roles · 75 employers) |
|---|---|---|
| Open roles in this specialty | 3 | 517 |
| Open roles in the wider Software, Firmware & Systems family | 8 | 5925 · 143 employers |
| Median days open | 29 d | 69 d (−40 d vs this employer) |
| Median salary (USD postings) | — | $232k |
Skills observed across this category: generative AI inference, performance modeling, ML workloads, distributed ML execution, system architecture, C++
Who's hiring in this category
- Qualcomm · 109 open roles · median 99 d
- NVIDIA · 87 open roles · median 70 d
- AMD · 44 open roles · median 62 d
- Micron Technology · 32 open roles · median 42 d
- Mobileye · 21 open roles · median 58 d
- Analog Devices · 14 open roles · median 27 d
How we counted: 517 open AI/ML Hardware Engineering (AI ML Hardware) roles from 75 employers tracked in the SemiconductorJobs index, counted 24 Sept 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.