DevOpsDays Rockies 2026

Why Most SLOs Fail
2026-09-22 , Main Stage

Most teams will tell you they have SLOs. Most of them are doomed to fail. Not because their metrics are bad or their tooling is wrong, but because reliability has no seat at the table. SLOs are useless unless they inform your roadmap. If error budget exhaustion doesn't change what gets built next sprint, you don't have an SLO program. You have an expensive monitoring system.

This talk draws on firsthand experience leading SRE initiatives and tech consulting to name the organizational patterns that make SLO programs fail, regardless of how much effort teams put in. You'll leave with a five-question diagnostic you can run at your next retro to find out exactly where your organization stands, and a clear picture of what it actually takes to give reliability the authority to change behavior.


This talk is drawn from direct SRE leadership and consulting work experience assessing SLO programs across engineering organizations of varying size and maturity. The failure modes described are ones I've observed repeatedly in the field- in organizations that had done everything "right" on paper and still saw no material improvement in production reliability.

Why this talk: Most SLO content focuses on the technical implementation- how to define SLIs, calculate error budgets, and build dashboards. That's not the problem. Teams that have done all of that are still failing. The gap is organizational, not technical, and it's rarely discussed.

Outline:

  • Why SLO programs fail even when teams do everything right
  • Four antipatterns observed firsthand that guarantee failure, regardless of tooling or effort
  • A maturity model for SLO programs- and why the jump from immature to minimal maturity isn't technical
  • What it practically takes to give reliability a seat at the table: exec commitment, enforced error budget policy, and product in the room
  • A five-question diagnostic practitioners can run immediately to assess where their organization stands

Audience takeaways:

  • The reason most SLO programs fail isn't technical- reliability lacks organizational authority, and no dashboard fixes that
  • A concrete maturity model to locate where your organization actually is, not where you think it is
  • Five diagnostic questions to run at your next retro that will tell you everything about your SLO program's health

Speaker background: ~20 years in tech, ~10 in DevOps/SRE. Seasoned public speaker.

Amin Astaneh is the founder of Certo Modo, a consultancy that helps technology companies strengthen reliability and operational practices. Drawing on two decades of experience scaling production systems of all sizes from scrappy startups to Big Tech companies, he brings both technical depth and hard-earned practical insight. He is also the host of the Reliability Rebels podcast, where he explores the social side of DevOps and Site Reliability Engineering. When he’s not helping teams, Amin is a digital nomad-often found camping out of his customized pickup truck and exploring new places.