Site Reliability Engineering Lead - Austin, Texas, United States
Originally posted 29 September 2026 by the employer.
About the role
This role is to build and lead a new Site Reliability Engineering (SRE) organization responsible for the production operation of a rapidly scaling AI supercomputing platform.
What you'll do
- Build and develop the SRE team from initial formation through full 24x7x365 production operations.
- Define the operating model for the SRE organization, including staffing, escalation paths, on-call responsibilities, and incident management.
- Embed with platform engineering teams to ensure reliability, serviceability, observability, and operational requirements are incorporated into the platform.
- Lead the development of SLOs, operational health indicators, alerting standards, and reliability reporting for large-scale production infrastructure.
- Lead or participate in major production incidents and establish a blameless post-incident review process.
- Ensure the SRE organization can rapidly diagnose and mitigate issues across compute, networking, storage, and supporting infrastructure.
What you'll need
- Significant experience leading or building an SRE, Production Engineering, or Infrastructure Reliability function supporting large-scale, highly available production infrastructure.
- Experience taking a new or rapidly evolving platform through production readiness, launch, stabilization, and ongoing operation.
- Experience building and operating sustainable 24x7x365 production support or on-call organizations.
- Strong understanding of modern Site Reliability Engineering principles, including SLOs, incident management, observability, automation, toil reduction, capacity management, and production readiness.
- Strong incident leadership experience, including managing high-severity, multi-team production incidents.
- A strong systems engineering background with a working understanding of Linux, networking, storage, distributed systems, automation, and production infrastructure.
Nice to have
- High-performance networking technologies such as InfiniBand, RDMA, RoCE, or large-scale Ethernet fabrics.
- Large-scale parallel or distributed storage systems.
- Workload schedulers and orchestration platforms such as Kubernetes, Slurm, or comparable systems.
- Production observability and service-level monitoring for complex distributed infrastructure.
- Custom or early-generation hardware, firmware, or environments where hardware and software are developed concurrently.
- Operational relationships and escalation processes involving datacenter operations and infrastructure/hardware/networking vendors.
- Rapid capacity expansion, datacenter migration, or transitions between temporary and permanent production environments.
Skills: SRE organization, production operations, reliability engineering, incident management, automation, distributed systems
This role has been open 11 days — well below the 52-day median for Infrastructure/Platform Software roles.
Infrastructure/Platform Software · Infrastructure Platform
|
Open roles in category
931
|
Median days open
52 d
|
Median salary
$227k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | Graphcore | All employers we track in this specialty (931 roles · 86 employers) |
|---|---|---|
| Open roles in this specialty | 44 | 931 |
| Open roles in the wider Software, Firmware & Systems family | 91 | 6035 · 146 employers |
| Median days open | 86 d | 52 d (+34 d vs this employer) |
| Median salary (USD postings) | — | $227k |
Skills observed across this category: SRE organization, production operations, reliability engineering, incident management, automation, distributed systems
Who's hiring in this category
- NVIDIA · 260 open roles · median 47 d
- Qualcomm · 61 open roles · median 57 d
- AMD · 52 open roles · median 29 d
- Graphcore (this employer) · 44 open roles · median 85 d
- Cerebras · 39 open roles · median 75 d
- Intel Corporation · 36 open roles · median 18 d
How we counted: 931 open Infrastructure/Platform Software (Infrastructure Platform) roles from 86 employers tracked in the SemiconductorJobs index, counted 10 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.