Skip to main content
Ampere Computing

AI Accelerator, Software Engineer- Graph Optimization/Compilers - Santa Clara, CA, United States

Ampere Computing Remote friendly (Santa Clara, CA, United States) 18 days ago
Embedded & Systems Software

$159,000 - $239,000 USD yearly

This role focuses on optimizing deep learning computational graphs for Ampere AI accelerators.

About the role

This role involves optimizing deep learning computational graphs to maximize the performance, efficiency, and scalability of Ampere's AI accelerator hardware. You will work across the software stack, from model frameworks and inference-serving systems to graph optimization, compiler infrastructure, runtimes, and compute kernels.

What you'll do

- Optimize computational graphs for performance, throughput, latency, memory efficiency, and power efficiency on Ampere AI accelerators. - Enable and optimize models, frameworks, and inference platforms, including PyTorch, Llama.cpp, vLLM, and SGLang. - Develop graph-level optimizations such as operator fusion, pattern matching, redundancy elimination, constant folding, layout optimization, memory planning, quantization, and accelerator offload. - Optimize transformer and LLM workloads, including dynamic shapes, attention mechanisms, KV-cache management, and mixed-precision execution. - Analyze end-to-end performance across frameworks, compilers, runtimes, kernels, and hardware. - Build profiling, benchmarking, validation, and performance-regression infrastructure.

What you'll need

- Bachelor’s degree in Computer Science, Computer Engineering, Mathematics, or a related technical field & 5 years of relevant experience; or a Master’s degree with & 3 years of relevant experience. - Strong foundations in algorithms, data structures, graph algorithms, computational complexity, and systems programming. - Proficiency in Python and C/C++. - Strong ability to reason about execution dependencies, memory movement, numerical correctness, and hardware execution behavior. - Familiarity with deep learning concepts, neural-network architectures, tensor operations, numerical precision, quantization, and memory layout.

Nice to have

- Experience diagnosing performance issues through profiling, benchmarking, tracing, or hardware-level analysis. - Experience with CUDA, ROCm, OpenCL, SYCL, Triton, GPU programming, NPU programming, or other accelerator architectures. - Familiarity with transformer models, LLM inference, attention mechanisms, KV-cache optimization, speculative decoding, mixed-precision execution, or sparsity. - Demonstrated exceptional problem-solving ability—IOI medal, ACM ICPC medal, Codeforces Grandmaster, USACO Platinum, or equivalent achievement in research or production engineering. - Experience using AI-assisted development tools.

Skills: AI Accelerator, Graph Optimization, Compilers, deep learning computational graphs, LLM workloads, inference platforms

Market context

This role has been open 7 days — well below the 41-day median for Infrastructure/Platform Software roles.

Infrastructure/Platform Software · AI ML Hardware

Open roles in category
123
Median days open
41 d
Median salary
$236k
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
How Ampere Computing compares in Infrastructure/Platform Software hiring
Metric Ampere Computing All employers we track in this specialty (123 roles · 29 employers)
Open roles in this specialty 2 123
Open roles in the wider Software, Firmware & Systems family 25 5791 · 144 employers
Median days open 14 d 41 d (−27 d vs this employer)
Median salary (USD postings) — $236k

Skills observed across this category: AI Accelerator, Graph Optimization, Compilers, deep learning computational graphs, LLM workloads, inference platforms

Who's hiring in this category

  • NVIDIA · 28 open roles · median 23 d
  • Qualcomm · 22 open roles · median 26 d
  • AMD · 12 open roles · median 25 d
  • Graphcore · 9 open roles · median 72 d
  • NXP Semiconductors · 7 open roles · median 86 d
  • Intel Corporation · 5 open roles · median 21 d

How we counted: 123 open Infrastructure/Platform Software (AI ML Hardware) roles from 29 employers tracked in the SemiconductorJobs index, counted 17 Sept 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.

Apply now
Santa Clara, CA, United States
Hybrid
$159,000 - $239,000 USD yearly
18 days ago

Share this job