Originally posted 21 August 2026 by the employer.
This role focuses on post-training quantization methods for large language models (LLMs) and other ML applications for optical inference engines.
About the role
This role involves developing advanced post-training quantization methods for large language models (LLMs), diffusion models, and other ML applications.
What you'll do
- Develop and execute hardware-aware post-training methods for full model quantization.
- Investigate preconditioning and formulate quantization as non-convex, discrete, constrained, or second-order optimization.
- Design controlled numerical experiments to understand potential improvements and secondary effects due to analog processing hardware.
- Build research-quality implementations and reproducible experiment harnesses for testing candidate methods.
- Adapt models from open-source repositories and customer private models, including PyTorch, Triton, JAX, and emerging frameworks.
- Design and execute re-quantization, retraining, and other model adaptation techniques to minimize accuracy loss during precision reduction.
What you'll need
- PhD, or equivalent research experience, in machine learning, applied mathematics, optimization, numerical analysis, or computer science.
- 5+ years of experience in machine learning, with at least 3 years focused on model optimization and deployment.
- Research or advanced engineering experience in neural network quantization, model compression, numerical optimization, or efficient inference.
- Strong knowledge of numerical linear algebra, including matrix factorizations, conditioning, covariance estimation, and iterative methods.
- Experience with one or more of non-convex optimization, discrete optimization, manifold optimization, second-order methods, or constrained optimization.
- Strong proficiency in PyTorch and familiarity with other ML frameworks, including JAX, Triton, and TensorFlow.
Nice to have
- Experience with low-precision inference optimization (INT8, FP8, or lower).
- Background in analog or optical computing architectures.
- Knowledge of in-memory computing paradigms and matrix-vector multiplication acceleration.
- Knowledge of randomized numerical linear algebra, sketching, or structured transforms.
- Publications in quantization, optimization, numerical linear algebra, model compression, or efficient ML.
- Experience with vector quantization, lattice methods, learned codebooks, or rate-distortion ideas.
- Experience with large-scale batch inference optimization.
- Familiarity with prefill versus decode optimization strategies in LLM inference.
- Experience conducting experiments on models large enough to expose scaling and generalization problems.
Skills: post-training quantization, LLMs, optical inference engines, PyTorch, model optimization
This role has been open 46 days — well below the 76-day median for AI/ML Hardware Engineering roles.
AI/ML Hardware Engineering · AI ML Hardware
|
Open roles in category
504
|
Median days open
76 d
|
Median salary
$230k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | Neurophos | All employers we track in this specialty (504 roles · 75 employers) |
|---|---|---|
| Open roles in this specialty | 4 | 504 |
| Open roles in the wider Software, Firmware & Systems family | 9 | 5996 · 146 employers |
| Median days open | 32 d | 76 d (−44 d vs this employer) |
| Median salary (USD postings) | — | $230k |
Skills observed across this category: post-training quantization, LLMs, optical inference engines, PyTorch, model optimization
Who's hiring in this category
- Qualcomm · 99 open roles · median 109 d
- NVIDIA · 84 open roles · median 72 d
- AMD · 42 open roles · median 66 d
- Micron Technology · 30 open roles · median 58 d
- Mobileye · 18 open roles · median 73 d
- NXP Semiconductors · 15 open roles · median 63 d
How we counted: 504 open AI/ML Hardware Engineering (AI ML Hardware) roles from 75 employers tracked in the SemiconductorJobs index, counted 6 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.