Skip to main content
d-Matrix

Sr Staff Site Reliability Engineer, AI Infrastructure - Santa Clara, United States

d-Matrix Santa Clara, United States Full-time 11 days ago
Embedded & Systems Software

Originally posted 29 September 2026 by the employer.

About the role

This role is for a Sr Staff Site Reliability Engineer focused on AI Infrastructure, responsible for the infrastructure layer relied upon by engineering teams and customers, including colocation facilities, on-premises GPU clusters, cloud environments, and platform services. This position is a core member of the SRE team, driving reliability, automation, and observability across colo, on-premises lab, and cloud environments, supporting CI/CD, QA, HPC workloads for silicon development, and customer-facing deployments.

What you'll do

  • Own reliability and availability across colo server fleets, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.
  • Perform hands-on infrastructure work, including server provisioning, OS configuration, networking, storage, and hardware troubleshooting from bare metal through auto-scaling Kubernetes environments.
  • Drive provisioning, deployment, and operational changes through Terraform and/or Ansible, contributing to shared IaC modules.
  • Build and document automation for host lifecycle management, fleet health checks, auto-remediation, self-service tooling, and networking automation for cluster interconnects and lab configurations.
  • Design and maintain monitoring, alerting, and SLIs using Prometheus/Grafana, DataDog, Splunk, or equivalent, contributing to AIOps-driven detection workflows.
  • Participate in on-call rotation, triaging and resolving incidents from bare metal to application layer, and produce RCAs for P0/P1 incidents.

What you'll need

  • Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience); 7+ years in SRE, infrastructure engineering, or systems administration.
  • Strong Linux systems knowledge with hands-on colocation or on-premises server infrastructure, including networking, storage, systemd, kernel parameters, performance diagnostics, physical hardware, rack networking, and bare-metal provisioning.
  • Production IaC experience with Terraform and/or Ansible, including writing and maintaining configurations.
  • Kubernetes operational experience: cluster troubleshooting, workload management, storage, and networking.
  • Experience with observability tooling: Prometheus/Grafana, DataDog, Splunk, or equivalent, including building dashboards and writing alert rules.
  • Production-quality Python and/or Bash scripting, paired with incident response experience: structured triage, RCA production, and follow-through on action items.

Nice to have

  • Experience operating customer-facing infrastructure or platform services with external reliability expectations.
  • Cloud infrastructure operations across AWS, Azure, or GCP, including hybrid environments spanning cloud and on-prem.
  • Experience deploying and operating AI-driven infrastructure tools: AIOps platforms, intelligent alerting, anomaly detection, or LLM-assisted diagnostics in production.
  • HPC job scheduler experience: Slurm, LSF, or equivalent.
  • Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink.
  • Experience with large-scale infrastructure automation: host lifecycle management, fleet auto-healing, or AIOps-driven operations, building tooling that reduces manual intervention.

Skills: Kubernetes, Terraform, Ansible, AI infrastructure, Prometheus/Grafana

Market context

This role has been open 9 days — well below the 52-day median for Infrastructure/Platform Software roles.

Infrastructure/Platform Software · Infrastructure Platform

Open roles in category
931
Median days open
52 d
Median salary
$227k
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
How d-Matrix compares in Infrastructure/Platform Software hiring
Metric d-Matrix All employers we track in this specialty (931 roles · 86 employers)
Open roles in this specialty 7 931
Open roles in the wider Software, Firmware & Systems family 18 6078 · 146 employers
Median days open 13 d 52 d (−39 d vs this employer)
Median salary (USD postings) — $227k

Skills observed across this category: Kubernetes, Terraform, Ansible, AI infrastructure, Prometheus/Grafana

Who's hiring in this category

  • NVIDIA · 260 open roles · median 47 d
  • Qualcomm · 61 open roles · median 57 d
  • AMD · 52 open roles · median 29 d
  • Graphcore · 44 open roles · median 85 d
  • Cerebras · 39 open roles · median 75 d
  • Intel Corporation · 36 open roles · median 18 d

How we counted: 931 open Infrastructure/Platform Software (Infrastructure Platform) roles from 86 employers tracked in the SemiconductorJobs index, counted 10 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.

Apply now
Santa Clara, United States
On-site
Full-time
11 days ago

Share this job