Sr AI/ML Performance Engineer - Remote
Originally posted 10 September 2026 by the employer — open 9 days.
About the role
What you'll do
- Drive AI/ML ASIC architecture integrating AI Storage with GPU/TPU/xPU accelerators, focusing on I/O subsystems connected over UCIe/PCIe/CXL.
- Author architecture specifications for AI/ML xPU based Accelerators using AI Storage Solutions.
- Define I/O subsystem and PCIe DMA architectures, including interactions with embedded processor-subsystems, Network on Chip, and Memory controllers.
- Create flexible and modular I/O subsystem architectures deployable in Chiplet, monolithic, or 3D form factors.
- Analyze LLM workloads and characterize ASIC and competitive datacenter/AI solutions to identify performance improvement opportunities.
- Architect memory-efficient inference/training systems using techniques like pruning, quantization with MX format, continuous batching/chunked prefill, and speculative decoding.
What you'll need
- Bachelors or Masters in Computer/Electrical Engineering with 1-3 years of architecture experience authoring specifications.
- Technical background architecting ASIC, SoC, or I/O subsystems involving PCIe/UCIe/CXL and DMA engines.
- Knowledge of I/O Subsystem and DMA interactions with internal embedded processor-subsystems (x86, RISC-V or ARM) and external host CPU.
- Understanding of computer/graphics architecture, ML, and LLM.
- Experience architecting GPU/TPU/xPU Accelerator systems with optimized high bandwidth memory hierarchy and frontend architecture for multi-trillion parameter LLM training/inference.
- Deep experience optimizing large-scale ML systems and GPU architectures.
Nice to have
- Familiarity with UCIe, CXL, NVLink, or UAL microarchitecture and protocols.
- Familiarity with High-speed networking: InfiniBand, RDMA, NVLink.
- Expert knowledge of transformer architectures, attention mechanisms, and model parallelism techniques.
- Multi-disciplinary experience, including familiarity with Firmware and ASIC design.
- Expertise in CUDA programming, GPU memory hierarchies, and hardware-specific optimizations.
- Proven track record architecting distributed training systems handling large scale systems.
- Previous experience with NVMe storage systems, protocols, and NAND flash.
Skills: AI/ML ASIC architecture, GPU/TPU/xPU accelerators, UCIe/PCIe/CXL, chiplet, LLM workload analysis, HBM
This role has been open 9 days — well below the 66-day median for AI/ML Hardware Engineering roles.
AI/ML Hardware Engineering · AI ML Hardware
|
Open roles in category
517
|
Median days open
66 d
|
Median salary
$232k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | SanDisk | All employers we track in this specialty (517 roles · 76 employers) |
|---|---|---|
| Open roles in this specialty | 8 | 517 |
| Open roles in the wider Software, Firmware & Systems family | 60 | 5787 · 143 employers |
| Median days open | 47 d | 66 d (−19 d vs this employer) |
| Median salary (USD postings) | — | $232k |
Skills observed across this category: AI/ML ASIC architecture, GPU/TPU/xPU accelerators, UCIe/PCIe/CXL, chiplet, LLM workload analysis, HBM
Who's hiring in this category
- Qualcomm · 106 open roles · median 98 d
- NVIDIA · 85 open roles · median 66 d
- AMD · 43 open roles · median 61 d
- Micron Technology · 32 open roles · median 38 d
- Mobileye · 25 open roles · median 52 d
- Analog Devices · 14 open roles · median 23 d
How we counted: 517 open AI/ML Hardware Engineering (AI ML Hardware) roles from 76 employers tracked in the SemiconductorJobs index, counted 20 Sept 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.