Beer & Servers Don't Mix

The Biggest Problem With Our Codebase Is the Developers

Or: What Platform Teams Actually Exist For (And Why Most Get It Completely Backwards)

It was a Tuesday afternoon on level 6 when someone asked the question that ended the meeting. Not loudly — just dropped it into a lull in conversation, the way you drop something heavy onto vinyl flooring. The question was: “How many different ways do we handle authentication token expiry across our services?” The whiteboard still had yesterday’s architecture diagram on it. The Thai iced teas had gone warm. Someone started counting on their fingers, then ran out of fingers. The answer, after a few more minutes of uncomfortable archaeology through Slack and Confluence, was eleven. Three of them were wrong. We wouldn’t find out which three for another six months.

That wasn’t a hiring problem. That wasn’t a code review problem. That was a scale problem dressed up as a technical one — and the solution wasn’t going to come from a retrospective or a stronger linting rule.

“Hell is other people.”* — Jean-Paul Sartre*

Sartre was writing about existential conflict. He was also, without knowing it, describing what happens when 200 engineers solve the same problem independently.

Here’s the uncomfortable truth that platform teams exist to address: developers — given complete freedom — will make individually rational choices that collectively produce chaos. Not because they’re careless. Not because they’re bad engineers. Because local optimisation doesn’t equal global optimisation. Every engineer who solves the A/B testing problem their own way is making a sensible decision. The result across a thousand engineers is five different implementations, three of which are buggy, none of which interoperate, and all of which need to be maintained indefinitely by people who weren’t there when they were written.

The central tension of engineering at scale is this: freedom is the source of developer joy, and the source of organisational entropy. Platform teams exist to manage that tension — not by eliminating freedom, but by making the right path the easy path.

The guardrail isn’t a cage. It’s what lets you drive fast.

What a Platform Actually Is (And Isn’t)

Most people hear “platform team” and think: Kubernetes. CI/CD pipelines. Cloud infrastructure. That’s one kind of platform — the runtime kind. But the more interesting and underappreciated kind is the application platform: the shared libraries, middleware, conventions, and tooling that shape how engineers write code, not just where that code runs.

Think about what happens in a large engineering organisation without one. Four separate teams need client-side load balancing. Hardware load balancers don’t work at the scale you’re operating — too expensive, and the heartbeat mechanism means a dead node causes a segment of traffic to fail for minutes before the decision to remove it propagates. So each team builds their own implementation. Not because they’re careless. Because the organisation hasn’t yet built the thing that should exist to solve this once. Each team locally optimised. The collective result was four implementations, fragmented across the codebase, some of them sharing organically across team boundaries, others silently diverging. Eventually a platform team was formed for this space and the consolidation began — slowly, as migrations always do.

This is what scale looks like without a platform. Not dramatic failure. Just quiet, compounding fragmentation.

“The strength of the team is each individual member. The strength of each member is the team.”* — Phil Jackson*

Jackson was talking about the Chicago Bulls. He was also, unknowingly, describing the economics of shared platform infrastructure. The individual engineer is more capable when the team around them has already solved the foundational problems. The platform is the mechanism by which that happens at scale.

The test for whether something belongs in the platform is simple: would at least five teams build this independently if it didn’t exist? If yes, the platform should own it. If no, you might be building premature abstraction.

The Scale Threshold

Platform teams make no sense below a certain size. Two engineers don’t need a platform — they need a conversation. At ten engineers, shared documentation is probably sufficient. The investment starts paying off around eight to ten teams, and the ROI compounds from there.

Below that threshold, organic sharing works. Engineers naturally cross-pollinate solutions, grab code from each other’s repositories, have corridor conversations. Joel Spolsky once wrote about “the Joel Test” as a quick measure of engineering hygiene — platform investment is something like that, but for organisations rather than teams. You can operate without it for a while. The symptoms are subtle at first.

Above the threshold, organic collaboration breaks down faster than documentation can keep up. The divergence tax compounds. You stop being able to answer basic questions — “how many ways do we do X?” — without an audit. New engineers joining take months to understand “how we do things here” because “how we do things here” is now twelve different answers depending on which team you landed on.

This connects directly to the paved paths work we’ve covered in this series — the platform is the delivery mechanism for paved paths at scale. The path is the convention. The platform is the infrastructure that makes following the path easier than not following it.

The Modular Lesson (That Microsoft Already Learned)

The original .NET Framework gave you everything. The problem was you took all of it whether you needed it or not — the kitchen sink came standard. One of the central design decisions in .NET Core (later .NET 5+) was a shift to modular, independent libraries with inversion of control at the seams. Don’t want the full HTTP stack? Don’t include it. Need to swap the logging abstraction? Write a new implementation for the interface everything is using already and override it. The framework became composable.

This is the design philosophy platform teams should steal — not the technology, the philosophy. Build modular. Define clear interfaces. Inversion of control at every seam so consumers can override what they need to override. The platform should be composable, not monolithic.

We learned this lesson at Agoda the expensive way.

When breaking up a large monolith, the business domain separation was done well. But the cross-cutting concerns — auth, observability, HTTP pipeline middleware, attribution logic — were bundled into a single shared library. The intent was good: “everything you need, packaged up nicely.” The result was what Joe Armstrong, creator of Erlang, once described as the Gorilla/Banana problem: you wanted a banana, you got a gorilla holding a banana, and the entire jungle along with it.

One initialisation method: AddPlatform(). Inside it, anything could go wrong. When it did, nobody knew how it was woven together.

Two failure modes followed immediately. The first was a debugging black box — “why can’t I debug this locally?” became a constant escalation. When AddPlatform() failed, engineers had no visibility into what was happening inside it. The second was structural: because the HTTP pipeline lived inside the platform library, teams needing to make changes to cross-cutting concerns — attribution logic being the example — had to go through the platform team. Not because the platform team was obstructive. Because the architecture made them the mandatory path. A business logic change became a platform ticket.

The platform team didn’t become a bottleneck because they were slow or territorial. They became one because the design created the dependency.

Microsoft had already figured this out. We learned it again, from scratch, the usual way.

Divergence Is a Signal, Not a Problem

Here’s the argument that most platform teams get exactly backwards.

When an engineer goes off the platform path — builds something the platform doesn’t support, rolls their own solution — the instinct is to treat it as a problem to be standardised away. Get everyone back on the path. Raise it in the next architecture review. File a ticket to add it to the roadmap.

This is exactly wrong.

Divergence is innovation. The engineer who went off-path did so because the path didn’t solve their problem. That’s not insubordination — it’s a signal. They found a gap in the platform. And they’re the best possible person to have found it, because they’re closest to the problem it failed to address.

The right model: track divergence and feed it back into the platform roadmap. Not every divergence becomes a platform feature — some are genuinely team-specific. But the pattern of where teams diverge tells you exactly where the platform is under-invested. It’s free product research. It’s the thing most platform teams would pay for if they could, sitting right there in their own codebase, being treated as a compliance problem.

Practically, this means building the platform with explicit extension points — seams where teams expect to plug in their own behaviour. It means creating a lightweight channel for teams to contribute divergent solutions back into the platform, rather than treating every off-path decision as a governance failure. Martin Fowler describes this as weak code ownership applied at scale — modules have owners, but the contribution model isn’t closed.

The framework: platforms that prevent divergence become bottlenecks. Platforms that channel divergence become better platforms.

The Guardrail vs. The Gate

A guardrail prevents bad outcomes without preventing movement. You can still drive fast. You can still choose your route. The guardrail activates when you’re about to go off a cliff.

A gate stops you. Full stop. Someone decides whether you may proceed.

Most platform teams start as guardrails and gradually become gates. The progression is always the same: build shared library → developers use it → edge case emerges → platform team adds a constraint to prevent misuse → another edge case → another constraint → the library now has seventeen configuration options, a ticket process for exceptions, and an approval workflow for new adopters.

You’ve built a gate and called it a platform.

The guardrail version has sensible defaults that cover 90% of cases. For the other 10%, you can override. You don’t need to ask permission. The guardrail is there for the cases where developers are about to make a decision they’ll regret — not for every decision.

There’s a simple test for which one you’ve built: if using the platform increases a developer’s cognitive load compared to rolling their own solution, the platform has failed. The whole point is to reduce cognitive load — to make the right thing the easy thing, the default thing. A developer who avoids the platform because it’s more work than the alternative is a developer the platform has lost. And a developer the platform has lost will write their own implementation, and you’re back to eleven ways of handling token expiry.

Team Topologies puts it clearly: the measure of a platform team’s success is how little stream-aligned teams need to think about it. Not how many features the platform has. Not how comprehensive the documentation is. How little friction it creates.

Platform as Product (The Mindset Shift)

The most underappreciated insight in platform engineering: your developers are your customers. Every instinct you have about product development — understand the user’s job to be done, measure adoption, gather feedback, reduce friction in the onboarding experience — applies here.

Most platform teams operate as internal utilities. They build what they think the organisation needs. They respond to requests. They define roadmaps based on their own assessment of the architecture. This works until it doesn’t — until adoption stalls because the platform solves the wrong problems, until teams work around it rather than with it, until the platform team is busy maintaining things nobody uses.

One quarter, our platform team took on a DX initiative: move local development environments from shared QA servers to local test containers. The motivation was real — if another team deployed a broken package to a shared QA server your developers depended on, your entire team could be blocked for hours waiting for a fix that wasn’t yours to make. Local isolation solved the dependency problem. But it introduced a new one: startup time. A freshly initialised local container is cold. A shared QA server is warm and pre-loaded. Solve one problem badly and you trade a fragile-but-fast environment for a reliable-but-slow one.

The platform team used local development metrics to measure startup time impact and ensure it stayed within acceptable bounds. Not a subjective argument about whether the new setup “felt” slower — data on whether it actually was, and by how much. That’s what allowed the migration to proceed with confidence rather than stalling in a preference debate.

This connects directly to the inner loop work we’ve covered — developer experience metrics aren’t just useful for measuring problems. They’re the mechanism by which platform teams are held accountable for the quality of the tools they build.

The product mindset requires:

  • Adoption metrics. Not just “is the library available” but “are teams actually using it?” And if not, why not?- Office hours and pairing. Platform teams that sit in an ivory tower and throw APIs over the wall produce platforms nobody loves. Platform teams that work alongside consuming teams build things that fit.- Explicit deprecation processes. Nothing erodes trust faster than a breaking change with a short notice window.- Developer satisfaction as a first-class outcome. Accelerate shows clearly that deployment frequency and lead time for changes are correlated with engineering culture — the platform is a direct input into both.

Platform Teams Must Own a Production System

This is the mechanism that makes the product mindset actually work, and it deserves naming as a structural policy rather than a vague principle.

Platform teams that don’t own a system that runs on their own platform gradually lose the ability to feel what they’ve built. The setup docs get longer because nobody on the platform team had to follow them recently. The initialisation ceremony grows because nobody was inconvenienced by the complexity. The error messages stay cryptic because nobody on the platform team debugged them under pressure at 5pm.

We had an example of this done right. A platform engineer, working on the team’s own production system, noticed inconsistent error handling patterns across the codebase — different teams handling errors differently, no standardisation, no obvious right answer in the existing platform. Because he was using the platform himself rather than just building it, he saw the gap firsthand rather than hearing about it via ticket. He built an npm package to standardise error handling. Critically: because he’d had to add it himself, there was no complex setup documentation. No onboarding guide. No “read this before you start.” Just: install the package, add one snippet to your page-level components, done. The solution was shaped by the experience of being the first person to use it.

That’s the feedback loop that bad platform teams don’t have. Dogfooding isn’t just a quality assurance mechanism. It’s an empathy mechanism. The engineer who has to add their own error handling library to their own system at 4pm on a Friday will write a better library than the one who specifies it from a design document.

“In theory, there is no difference between theory and practice. In practice, there is.”* — *old engineering adage

The fix is structural: the platform team must own and operate a production system that uses the platform. Not a demo environment. Not a reference implementation. A real system, serving real traffic, that the platform team is on-call for. This is dogfooding, but with teeth.

The “That’s the Platform Team’s Job” Anti-Pattern

Name this one explicitly, because it lives in every large organisation and almost nobody talks about it.

A developer identifies a problem — something genuinely useful that the platform doesn’t yet address. They raise it. The response: “That’s the platform team’s job. File a ticket.” The ticket enters the backlog. The developer’s team ships their product deadline, rolls their own solution, and now there are two implementations of the thing. The platform team eventually builds their version, which the developer’s team has no incentive to migrate to because their version already works.

Cost: duplicated effort, inconsistency, developer frustration, and a platform roadmap that doesn’t reflect what teams actually need.

The fix isn’t to tell developers to wait. It’s to build a model where developers can contribute to the platform with appropriate oversight — what Tim O’Reilly coined — and Martin Fowler champions — as the innersource model. Not “commit to the platform repo without review,” but “here is the process by which a team with a real problem can extend the platform and have that extension considered for inclusion.”

The caveat: not everything a developer wants to add belongs in the platform. The platform team’s job includes saying no to contributions that are too team-specific, too untested, or at odds with the architecture direction. But the platform team that says no to everything and makes developers file tickets for everything has confused governance with gatekeeping.

This is a structural sibling to the silo problem we’ve written about before — the innersource model is one of the few mechanisms that structurally forces cross-team collaboration without requiring anyone to reorganise.

The Bottom Line

Platform teams exist because scale breaks the informal coordination mechanisms that work perfectly well at twenty engineers and catastrophically badly at two hundred. They exist not because developers can’t be trusted, but because developers can be trusted to solve every problem in front of them — including ones they shouldn’t each have to solve individually.

The platform’s job is to make the right path the easy path. To absorb the solved problems so teams can focus on the unsolved ones. To reduce the cognitive load of building a new service from “learn twelve different conventions” to “follow this template, extend where you need to.”

The platform teams that fail treat this as an infrastructure problem — something to govern, constrain, standardise. The platform teams that succeed treat it as a product problem — something to design, measure, iterate, and ultimately make their customers love.

“You can design and create, and build the most wonderful place in the world. But it takes people to make the dream a reality.”* — Walt Disney*

Your engineers are the people who make the architecture real. The platform exists to give them the best possible conditions to do that. If you’re building a gate and calling it a guardrail, you’re not protecting the codebase — you’re just adding a tollbooth to the road.

As for me — I’m going to go check how many different ways we currently handle retry logic across our services. I have a feeling the answer is going to require more than one hand.