Skip to main content
Tensordyne

Forward Deployed Inference Engineer - Munich, Germany

Tensordyne Munich, Germany 1 day ago
AI/ML Hardware

Originally posted 5 October 2026 by the employer.

About the role

This role involves turning customer AI workloads and model requests into optimized technical results running on Tensordyne hardware and software.

What you'll do

  • Turn customer workloads into fast, credible performance answers through profiling and benchmarking, defining KPIs, and comparing against baselines.
  • Convert and bring up customer models on the Tensordyne stack, validate numerical quality, identify bottlenecks, and work with compiler, runtime, kernel, and system teams.
  • Work directly with customers and partners on technical PoCs, integration, deployment, and debugging, translating requirements into measurable acceptance criteria.
  • Track profiling-to-hardware accuracy, explain material gaps, and flag missing capabilities in the compiler, SDK, inference server, or KV-cache management.
  • Turn repeated customer-specific learnings into reusable tooling, documentation, benchmarks, or product improvements.

What you'll need

  • Strong hands-on experience with AI models and inference systems, especially dense and MoE LLMs (Llama, DeepSeek, Qwen, GPT-OSS, Kimi, GLM), and VLM, speech and diffusion models.
  • Strong Python and PyTorch skills and the ability to understand and modify model code.
  • Experience profiling, benchmarking, or optimizing model inference and reasoning about latency, throughput, memory, and utilization.
  • Proficiency in using AI-powered developer tools (e.g., Claude Code, Cursor).

Nice to have

  • Experience with LLM serving and deployment stacks such as vLLM, SGLang or similar systems.
  • Experience working directly with customers or external technical partners.
  • Experience bringing models up on new accelerators or non-standard hardware, including performance debugging across framework/runtime/hardware boundaries.
  • Practical experience with production inference techniques or environments such as quantization, distributed inference, or Kubernetes.
  • Experience navigating and contributing to Rust codebases.

Skills: AI models, inference systems, LLMs, PyTorch, accelerators

Market context

This role has been open 0 days — well below the 76-day median for AI/ML Hardware Engineering roles.

AI/ML Hardware Engineering · AI ML Hardware

Open roles in category
504
Median days open
76 d
Median salary
$230k
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
How Tensordyne compares in AI/ML Hardware Engineering hiring
Metric Tensordyne All employers we track in this specialty (504 roles · 75 employers)
Open roles in this specialty 5 504
Open roles in the wider Software, Firmware & Systems family 10 5996 · 146 employers
Median days open 12 d 76 d (−64 d vs this employer)
Median salary (USD postings) — $230k

Skills observed across this category: AI models, inference systems, LLMs, PyTorch, accelerators

Who's hiring in this category

  • Qualcomm · 99 open roles · median 109 d
  • NVIDIA · 84 open roles · median 72 d
  • AMD · 42 open roles · median 66 d
  • Micron Technology · 30 open roles · median 58 d
  • Mobileye · 18 open roles · median 73 d
  • NXP Semiconductors · 15 open roles · median 63 d

How we counted: 504 open AI/ML Hardware Engineering (AI ML Hardware) roles from 75 employers tracked in the SemiconductorJobs index, counted 6 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.

Apply now
Munich, Germany
On-site
1 day ago

Share this job