DevOpsDays Rockies 2026

Brent Chapman

Brent Chapman helps companies prevent, prepare for, respond to, and learn from incidents in their IT services. Over more than three decades in Silicon Valley, he has designed, built, managed, and scaled IT infrastructure and teams for everything from embryonic startups to giants such as Google, Slack, and Atlassian. As a leader in Google's legendary SRE organization, he created the Incident Management at Google (IMAG) protocol now used throughout the company. Slack later recruited him to build and lead its incident management program, and he has since led and advised incident management efforts at many other companies large and small. Today he runs Great Circle Associates, providing incident management consulting and training, and he is writing the forthcoming book Incident Management for DevOps and SRE.

Brent brings a rare dual perspective from a parallel volunteer career in public safety emergency services: air search and rescue pilot and incident commander, emergency dispatcher and dispatch supervisor for Burning Man, and CERT instructor. He is coauthor of the highly regarded O'Reilly book Building Internet Firewalls, the creator of Majordomo (the classic open source Internet mailing list manager), and a frequent speaker at conferences worldwide.


Social Media

Blog: https://greatcircle.com/blog/
LinkedIn: https://www.linkedin.com/in/brentchapman
Mastodon: https://hachyderm.io/@brent_chapman


Session

09-22
09:15
60min
When Production Burns, You're the Fire Department (and the Fire Marshal, Too)
Brent Chapman

When something breaks in production, nobody calls 911; you and your team are the emergency responders. The fire service has spent more than a century learning how to organize people under pressure, and much of the tech industry has quietly adapted those lessons, often without realizing where they came from. But the fire department only covers half the job. Fire marshals cover everything outside the fire itself: preventing fires before they start, and investigating them afterward to figure out what really happened and what should change.

In this talk, we'll look at what DevOps and SRE teams can learn from both the fire department and the fire marshal: how the Incident Command System translates to managing technical incidents, and how the investigative mindset of the fire marshal can transform your incident reviews from blame-and-forget rituals into the most valuable learning your organization does. Expect practical lessons, a few war stories, and some surprising parallels between fighting fires and keeping services running.

Talks
Main Stage