advancedDevOps & CloudLive cohorts
SRE & Observability in Production
SLOs, error budgets and alerts that mean something
SRE & Observability in Production: SLOs, error budgets and alerts that mean something.
What you'll be able to do
- Define SLOs your team agrees on
- Instrument with OpenTelemetry
- Run a blameless incident review
Curriculum
Getting started
How this course workspreview
Setting up your environmentpreview13m
Starter repo & setup checklist
Core concepts
The mental model19m
Working through the happy path16m
What breaks in production (and why)
Tradeoffs and alternatives14m
Interactive walkthrough
Applied practice
Building the reference implementation27m
Guided exercise: build it yourself
Cheat sheet, templates & further reading
Build & assess
Knowledge check
SRE & Observability in Production: final assignment
Where to go next
Reviews
Sana Kapoor★★★★★
The burn-rate alerting section changed how our on-call works.