Originally posted 28 September 2026 by the employer.
This role focuses on mapping ML algorithms to the Cerebras architecture.
About the role
This role involves determining how ML algorithms should be mapped to the Cerebras architecture and evaluating their performance. The engineer will characterize the efficiency frontiers of emerging ML algorithms, spanning kernel-level and end-to-end performance.
What you'll do
- Build analytical and empirical performance models for ML training and inference algorithms.
- Characterize asymptotic behavior and identify how algorithmic trade-offs change with model size, sequence length, batch size, parallelism, and hardware scale.
- Construct Pareto frontiers across model quality, latency, throughput, memory footprint, communication, and compute cost.
- Develop prototype implementations and benchmarks for the Cerebras WSE and relevant GPU or software baselines.
- Analyze system behavior to identify kernel, compiler, runtime, communication, and algorithmic bottlenecks.
- Partner with researchers and kernel, compiler, runtime, inference, and architecture teams to recommend implementation and co-design directions.
What you'll need
- Bachelor’s, Master’s, PhD, or equivalent practical experience in Computer Science, Computer Engineering, Electrical Engineering, Mathematics, or a related field.
- Strong foundation in computer architecture, parallel computing, and systems performance.
- Strong understanding of machine learning fundamentals and ML systems, including how model and algorithmic choices affect compute, memory, communication, accuracy, and scaling behavior.
- Experience with analytical performance modeling, algorithmic complexity analysis, benchmarking, or system simulation.
- Proficiency in Python and comfort with C++.
- Experience profiling and debugging performance in an ML, HPC, CPU, GPU, or accelerator-based system.
Nice to have
- Experience with roofline analysis, CPU or GPU simulators, kernel optimization, or hardware–software co-design.
- Familiarity with CUDA, Triton, PyTorch, JAX, or open-source LLM training and inference systems.
- Understanding of transformer internals, including attention variants, KV-cache strategies, model parallelism, sparsity, quantization, and parallel generation.
- Research publications, patents, or significant open-source contributions related to ML systems, computer architecture, or computational efficiency.
- Experience evaluating technology choices for future hardware or software architectures.
Skills: ML algorithms, Cerebras architecture, performance modeling, AI accelerator, hardware-software co-design
This role has been open 1 day — well below the 71-day median for AI/ML Hardware Engineering roles.
AI/ML Hardware Engineering · AI ML Hardware
|
Open roles in category
501
|
Median days open
71 d
|
Median salary
$225k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | Cerebras | All employers we track in this specialty (501 roles · 76 employers) |
|---|---|---|
| Open roles in this specialty | 8 | 501 |
| Open roles in the wider Software, Firmware & Systems family | 76 | 5943 · 145 employers |
| Median days open | 65 d | 71 d (−6 d vs this employer) |
| Median salary (USD postings) | — | $225k |
Skills observed across this category: ML algorithms, Cerebras architecture, performance modeling, AI accelerator, hardware-software co-design
Who's hiring in this category
- Qualcomm · 106 open roles · median 104 d
- NVIDIA · 81 open roles · median 67 d
- AMD · 42 open roles · median 64 d
- Micron Technology · 29 open roles · median 54 d
- Mobileye · 20 open roles · median 67 d
- Analog Devices · 14 open roles · median 33 d
How we counted: 501 open AI/ML Hardware Engineering (AI ML Hardware) roles from 76 employers tracked in the SemiconductorJobs index, counted 30 Sept 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.