Originally posted 5 October 2026 by the employer.
About the role
This role involves turning customer AI workloads and model requests into optimized technical results running on Tensordyne hardware and software.
What you'll do
- Turn customer workloads into fast, credible performance answers through profiling and benchmarking, defining KPIs, and comparing against baselines.
- Convert and bring up customer models on the Tensordyne stack, validate numerical quality, identify bottlenecks, and work with compiler, runtime, kernel, and system teams.
- Work directly with customers and partners on technical PoCs, integration, deployment, and debugging, translating requirements into measurable acceptance criteria.
- Track profiling-to-hardware accuracy, explain material gaps, and flag missing capabilities in the compiler, SDK, inference server, or KV-cache management.
- Turn repeated customer-specific learnings into reusable tooling, documentation, benchmarks, or product improvements.
What you'll need
- Strong hands-on experience with AI models and inference systems, especially dense and MoE LLMs (Llama, DeepSeek, Qwen, GPT-OSS, Kimi, GLM), and VLM, speech and diffusion models.
- Strong Python and PyTorch skills and the ability to understand and modify model code.
- Experience profiling, benchmarking, or optimizing model inference and reasoning about latency, throughput, memory, and utilization.
- Proficiency in using AI-powered developer tools (e.g., Claude Code, Cursor).
Nice to have
- Experience with LLM serving and deployment stacks such as vLLM, SGLang or similar systems.
- Experience working directly with customers or external technical partners.
- Experience bringing models up on new accelerators or non-standard hardware, including performance debugging across framework/runtime/hardware boundaries.
- Practical experience with production inference techniques or environments such as quantization, distributed inference, or Kubernetes.
- Experience navigating and contributing to Rust codebases.
Skills: AI models, inference systems, LLMs, PyTorch, accelerators
This role has been open 0 days — well below the 76-day median for AI/ML Hardware Engineering roles.
AI/ML Hardware Engineering · AI ML Hardware
|
Open roles in category
504
|
Median days open
76 d
|
Median salary
$230k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | Tensordyne | All employers we track in this specialty (504 roles · 75 employers) |
|---|---|---|
| Open roles in this specialty | 5 | 504 |
| Open roles in the wider Software, Firmware & Systems family | 10 | 5996 · 146 employers |
| Median days open | 12 d | 76 d (−64 d vs this employer) |
| Median salary (USD postings) | — | $230k |
Skills observed across this category: AI models, inference systems, LLMs, PyTorch, accelerators
Who's hiring in this category
- Qualcomm · 99 open roles · median 109 d
- NVIDIA · 84 open roles · median 72 d
- AMD · 42 open roles · median 66 d
- Micron Technology · 30 open roles · median 58 d
- Mobileye · 18 open roles · median 73 d
- NXP Semiconductors · 15 open roles · median 63 d
How we counted: 504 open AI/ML Hardware Engineering (AI ML Hardware) roles from 75 employers tracked in the SemiconductorJobs index, counted 6 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.