Non-Engineering
Originally posted 23 September 2026 by the employer.
About the role
This role involves owning and modernizing the enterprise monitoring and observability program, shifting towards predictive detection and intelligent event correlation. The scope includes managing IT systems for availability and performance across tools such as Datadog, BigPanda, SolarWinds, Zabbix, and Grafana.
What you'll do
- Own and modernize the enterprise monitoring and observability program, focusing on predictive detection, automation, and intelligent event correlation.
- Manage specific IT systems for availability and performance, reporting anomalies and participating in evaluation of patches and new systems.
- Manage execution of support services for monitoring/observability tooling, adhering to service management processes and conducting root cause analysis.
- Build automated remediation and self-healing workflows, integrating monitoring data with ITSM/orchestration platforms.
- Partner with the cybersecurity/SOC team to surface observability data for security investigations.
- Plan and manage projects within established scope, timeline, budget, and quality guardrails.
What you'll need
- 6–10 years in IT operations, infrastructure engineering, or SRE.
- Hands-on experience with Datadog, BigPanda, SolarWinds, Zabbix, Grafana (or equivalents such as Prometheus, Splunk, AppDynamics).
- Demonstrated experience maturing an observability program, including architecture, predictive/anomaly-based alerting, and noise reduction.
- Strong scripting/automation ability (Python, PowerShell, or similar) and experience integrating monitoring platforms via APIs.
- Experience supporting a large-scale, global, hybrid (cloud + on-prem) enterprise environment.
- People management experience or readiness for a first-line management role with technical credibility.