DevOpsDays Barcelona 2026

Lessons from Building at Massive Scale

Operating systems at massive scale means designing for failure, traffic spikes, and constant change.

In this talk, I’ll share lessons from building and running distributed systems at companies like Canva and Slack, focusing on resilience, high availability, incident response, and SRE practices. Expect real-world examples of what breaks at scale—and how to design systems that survive it.


Building software at massive scale means designing systems that don’t just work—but continue to work under extreme load, unpredictable traffic spikes, and constant change.

In this talk, I’ll share lessons from building and operating distributed systems at companies like Canva and Slack, where platforms handle trillions of requests and support millions of users in real time. We’ll focus on what it actually takes to design for resilience and reliability in high-pressure environments.

We’ll cover topics such as:

  • Designing for high availability and graceful degradation
  • Handling traffic spikes and load asymmetry without cascading failures
  • Evolving architectures as scale and complexity grow
  • Real-world incident response: what breaks, why, and how to recover fast
  • Applying SRE practices to keep systems reliable at 24x7 global scale

This is a practical, experience-driven talk focused on the realities of operating systems at scale—where failure is inevitable, and resilience is designed.

Javier Turégano

Javier Turégano is a platform engineering leader with 20+ years of experience designing and operating large-scale distributed systems. He has held senior roles at Canva and Slack, where he worked on edge, gateway, identity, and service-to-service architectures handling trillions of requests and extreme traffic spikes.

His work focuses on system design, resilience, and SRE practices, including high-availability architectures, incident response, and operating reliable 24x7 platforms at global scale.

Now an independent consultant, Javier advises companies on scalable architecture and resilient system design, and runs engineering leadership bootcamps in Spain and Australia.

LinkedIn