The Hardest Platform Problems Weren't Technical
Platform Engineering is often presented as a problem of tooling: self-service platforms, golden paths, developer portals and automation. Those capabilities are important, but in my experience they are rarely the limiting factor.
At Cabify we found that most platform initiatives had a reasonable justification behind them. Reliability improvements, developer experience, compliance requirements, cost optimization and new capabilities all created value. The challenge was not deciding what was worth doing, but deciding what was worth doing first.
This talk shares lessons learned from running infrastructure and platform teams where the number of worthwhile opportunities consistently exceeded available capacity, and how that changed our approach and processes to Platform Engineering.
Platform Engineering is often presented as a problem of tooling. The discussion usually revolves around self-service capabilities, golden paths, developer portals, automation and developer experience. Those capabilities are important and we invest in many of them ourselves, but over the last few years I have come to believe that they are rarely the limiting factor.
In my experience, platform teams do not suffer from a lack of ideas. The opposite is usually true. Reliability improvements, developer experience initiatives, compliance requirements, cost optimization, technical debt reduction and new platform capabilities can all have a reasonable business case behind them. The difficulty is that they all compete for the same limited capacity.
At Cabify we have been trying to evolve from infrastructure as a collection of services towards Platform Engineering with a stronger product mindset. That journey has been useful because it helped make some of these tensions more visible. We found ourselves spending less time discussing how to implement things and more time discussing why we should implement them, what value they would create and what other work would be displaced as a consequence.
As the number of opportunities grew, we started experimenting with different approaches to prioritization. We introduced a single backlog for Platform, moved discussions upstream, explored Internal Product Advocates (IPAs) and tried to make value and expected outcomes more explicit. None of these approaches magically solved the problem, but they helped create better conversations around trade-offs and priorities.
In this talk I will share examples from infrastructure and platform work, including operational improvements, reliability initiatives, governance requirements and developer experience investments. The goal is not to present a prioritization framework or claim that we found the answer. Rather, I want to discuss a challenge that I suspect many platform teams face: how do you make decisions when most of the available options are worthwhile and there is never enough capacity to do all of them?
My hope is that attendees leave with a different perspective on Platform Engineering. Beyond the technology, the tooling and the automation, there is an ongoing exercise in deciding where to invest limited capacity. Those decisions shape the platform at least as much as any architectural choice, yet they receive far less attention in our industry discussions.
Abel Navarro is a Senior Engineering Manager working in Platform Engineering. During the last 25 years he has worked across infrastructure, cloud platforms, observability, databases, networking and developer experience in companies ranging from startups to large international organizations.
His current focus is helping platform teams operate as internal products, reducing cognitive load for engineers and improving the flow of value across the organization. Abel is particularly interested in the intersection between technology, organizational design and platform adoption.