SRE & Observability

Reliability as something you can measure: SLOs and error budgets, multi-burn-rate alerts, the four golden signals, incident severity and a definition of "recovered" that holds up. Plus observability on eBPF and OpenTelemetry — and traps like HTTP/2 on internal traffic that only surface under load.