Originally posted 29 September 2026 by the employer.
Build and maintain telemetry pipelines for large-scale, performance-critical production systems.
About the role
This role involves building and evolving systems for deep visibility into large-scale, performance-critical production systems. It focuses on developing and operating distributed systems that power AI inference.
What you'll do
- Design and implement observability instrumentation across services and platforms
- Build and maintain telemetry pipelines for metrics, logs, and traces at scale
- Develop internal observability platforms, libraries, and tooling
- Define and operationalize SLIs, SLOs, and alerting strategies
- Partner with engineers to make systems debuggable by design
- Create clear, actionable dashboards and alerts that reflect real system health
What you'll need
- Strong experience in backend or systems software engineering
- Proficiency in one or more of: Go, C++, Rust, Java, Python
- Solid understanding of distributed systems
- Solid understanding of networking fundamentals
- Solid understanding of concurrency and performance tradeoffs
- Hands-on experience with metrics, logs, and distributed tracing
- Hands-on experience with production monitoring and alerting
- Familiarity with tools such as OpenTelemetry, Prometheus, Grafana, Datadog, Elastic, Jaeger, Tempo
- Experience designing high-signal alerts
- Experience designing scalable telemetry pipelines
- Experience designing service-level indicators and objectives
Nice to have
- Experience in high-performance computing, AI/ML systems, or inference platforms
- Hardware-aware observability (accelerators, GPUs, custom hardware)
- Prior SRE or platform engineering background
- Experience debugging large-scale production incidents
- Building internal developer platforms or shared libraries
Skills: observability, distributed systems, telemetry pipelines, production monitoring, AI inference
This role has been open 11 days — well below the 52-day median for Infrastructure/Platform Software roles.
Infrastructure/Platform Software · Infrastructure Platform
|
Open roles in category
931
|
Median days open
52 d
|
Median salary
$227k
|
See the full market breakdown ▾Category comparison, skills in demand, and who else is hiring
| Metric | Cerebras | All employers we track in this specialty (931 roles · 86 employers) |
|---|---|---|
| Open roles in this specialty | 39 | 931 |
| Open roles in the wider Software, Firmware & Systems family | 77 | 6035 · 146 employers |
| Median days open | 76 d | 52 d (+24 d vs this employer) |
| Median salary (USD postings) | — | $227k |
Skills observed across this category: observability, distributed systems, telemetry pipelines, production monitoring, AI inference
Who's hiring in this category
- NVIDIA · 260 open roles · median 47 d
- Qualcomm · 61 open roles · median 57 d
- AMD · 52 open roles · median 29 d
- Graphcore · 44 open roles · median 85 d
- Cerebras (this employer) · 39 open roles · median 75 d
- Intel Corporation · 36 open roles · median 18 d
How we counted: 931 open Infrastructure/Platform Software (Infrastructure Platform) roles from 86 employers tracked in the SemiconductorJobs index, counted 10 Oct 2026. Specialty figures count only roles carrying this exact specialty label, so an employer's related work in neighbouring specialties is not included there — it is counted in the wider Software, Firmware & Systems family row. Figures refresh nightly.