Beer & Servers Don't Mix

Your New System Is Done. Now Good Luck Getting 40 Teams to Care.

Or: What Nobody Tells You About Retiring a Legacy System

Somchai had been quiet for about three seconds — long enough that someone had started talking about the next agenda item — when he looked up at the whiteboard and asked the question. The kind of pause that happens when someone is choosing their words carefully, not because they don’t know what to say, but because they know exactly what they’re about to do to the conversation. The whiteboard still had the migration plan on it in confident blue marker — arrows from the old system to the new one, client boxes lined up like dominoes ready to fall. The meeting had been going well. Everyone had agreed the new system was good. The plan made sense. And then Somchai said: “Have you considered just implementing the old system’s contracts in the new one so clients can change a URL and nothing else breaks?”

The whiteboard didn’t change. But the arrows on it suddenly looked different.

Let’s talk about the plan that nobody argues with in the planning meeting, but that quietly costs you years of your life anyway.

Here’s the shape of it. You build a new system. It’s good — maybe the best thing your team has shipped in years. Clean contracts, sensible design, the architecture you’d have built the first time if you’d known then what you know now. The plan: migrate all internal clients across. How hard can it be? You send out the announcement. You write the migration guide. You schedule the deprecation date.

And then you wait.

Some teams move fast. A few move in the first month, enthusiastic early adopters who actually read the announcement. Then the pace slows. Client F has a frozen deployment window until Q3. Client G’s team is two people down and won’t be able to prioritise this until after the rebrand. Client H’s lead just left, nobody’s quite sure who owns it now, and the deprecation notice is sitting unread in a team inbox that may or may not be actively monitored.

A year passes. You’re still supporting both systems.

I watched a friend live through this. A new API — well-designed, the right thing, clearly superior to what it replaced. The plan was standard playbook: build new thing, migrate clients, retire old thing. It took six years.

Six years of keeping the old system alive. Six years of patches, of incident reviews, of onboarding new engineers into the lore of a system that was supposed to be dead. By year four, my friend was having the same conversation on rotation — a gentle reminder email here, an escalation there, a meeting where someone who’d been cc’d on the original announcement said they weren’t sure this had ever been communicated to their team. In one of those meetings, I told him — cheekily, in the way I’d been saying it for years to nudge teams toward inner sourcing — “I accept pull requests.” The implication being: your migration, your pull request. His response: his team didn’t have the capacity to do the migration work across fifty client systems.

Let that sit for a moment. The team that created migration work across fifty client systems didn’t have capacity to do it themselves — so the plan was that fifty other teams would each find the time to do it for them. Nobody was the villain. Everyone was busy with their own priorities. That’s exactly the problem. You can’t build a migration plan whose success depends on fifty teams making your work their priority. That’s not a plan. It’s a wish.

The Maths Nobody Does at the Start

The instinct that drove the six-year migration is almost universal. The moment the new system is built, the team is mentally done. The old system is legacy — a liability. The plan is to kill it as fast as possible. So the default playbook is: announce, document, set a deprecation date, hope.

What almost nobody does is price the alternative up front.

The arithmetic isn’t complicated:

Option A — Hard cutover, clients own their migration:

  • Build the new system- Keep the old system alive while clients migrate — weeks, months, or years, depending on how many clients you have and how much they care- The ongoing support cost of the old system runs for exactly as long as the slowest client takes to move Option B — Build a compatibility surface in the new system:

  • Build the new system plus a layer that speaks the old system’s contracts- The old system can be decommissioned as soon as the new one has feature parity and compatibility — you’re no longer blocked on client timelines- Clients migrate quickly — it’s just a URL change; the question is whether implementing the old system’s contracts in the new one costs less than supporting the legacy system for however long that tail turns out to be without it. The crossover question is simple: is the cost of building the compatibility layer less than the cost of supporting the old system for the duration of the tail? If you’re migrating three internal clients you fully control, Option A is almost always right. If you’re migrating fifty clients with independent deployment windows, external vendors, and teams who will simply deprioritise your migration quarter after quarter — the maths turns against you fast.

The variable nobody prices honestly is migration fatigue. Chasing clients to update a config is not free. Every follow-up email, every escalation, every meeting where someone has to explain again why this is important — that cost is real and it lands on the team that’s doing the retiring. It doesn’t show up in the project estimate. It shows up in engineer morale, three years later, when the person responsible for the migration has moved on and the old system is still running because it’s load-bearing in ways nobody fully mapped.

The Pattern Has a Name

Eric Evans named it in Domain-Driven Design in 2003: the Anti-Corruption Layer. His original framing was defensive — when a new system must integrate with a legacy system whose model is messy, you build a translation layer to stop the mess leaking into your clean new design. It’s the architectural equivalent of a quarantine zone.

The migration variant inverts the direction. Instead of the new system protecting itself from the legacy, the new system mimics the legacy’s contracts so that its clients don’t have to change. The ACL faces outward, speaking the old language. The internals stay clean.

When Somchai raised this in our meeting, he didn’t call it an anti-corruption layer. He said: “implement the deprecated system’s contracts in the new system so clients can just change the URL.” The vocabulary came later when someone asked what this pattern was called. The plain English was what actually landed.

This is worth naming because there are people in your organisation right now who would recognise the description immediately and have never once connected it to the DDD term. The concept is intuitive once you hear it. The vocabulary is what makes it reach for-able — the difference between a tool you can grab when you need it and an insight you can only articulate after the fact.

It’s also worth distinguishing from the Strangler Fig pattern, because they’re constantly conflated. Martin Fowler’s Strangler Fig is about gradual replacement through routing — you proxy traffic from old to new, migrating feature by feature, until the old system is strangled out of existence. The ACL compatibility variant is different: the new system is already built. The question is whether its external contracts match what clients already speak. One is a routing mechanism. The other is a compatibility surface. They’re complementary, not synonymous.

Why the Room Goes Quiet

Here’s the thing about the moment Somchai asked his question. He owned one of the client systems. He was, effectively, one of the dominoes on that whiteboard.

He wasn’t an outside architect spotting a missed option in someone else’s plan. He was the person about to bear the migration cost, suggesting a smarter way to absorb it. What he was really asking was: before I go change my system, have you considered building a compatibility surface so I don’t have to?

The room went quiet because it forced everyone to ask a question the plan hadn’t asked: how long is this migration actually going to take, and who’s paying for the support cost while it happens?

The economics professor Thomas Sowell has a line that applies here with uncomfortable directness: “There are no solutions. There are only trade-offs.” The hard cutover isn’t free — it just hides the cost in the migration tail. The ACL isn’t free either — it hides the cost in the upfront build. The question is which cost you’d rather pay, and when. The plan that doesn’t ask the question is the plan that defaults to the hidden cost every time.

In this particular case, the ACL wasn’t built. Three clients — small enough that the hard cutover was the right call. The maths worked out the other way, which is exactly the point. The ACL isn’t always the answer. The point is to do the maths rather than default to the standard playbook without asking.

The Second-Order Risk Nobody Plans For

There’s a coda to this that I find darkly funny, because it happened.

One ACL implementation at Agoda was named — informally, in the code — the “shitification layer” by the engineers who built it. The name is accurate: it made the new system speak the old system’s language, and the old system’s language was not something anyone was proud of. The pattern worked. The old system was retired. Mission accomplished.

The shitification layer is still there.

The ACL is not a free lunch. If you build one, it is not self-cleaning. It needs an owner. It needs a sunset date that someone is accountable for. It needs to be treated as first-class infrastructure — not a convenient bridge you build and then forget about. The trap is that once you’ve absorbed the migration cost into the new system, the urgency to actually complete the migration drops to nearly zero. Clients have no pressure to move. The old contracts keep working. Months become years. Your clever architectural move becomes the next legacy system.

The surgeon and writer Atul Gawande wrote in The Checklist Manifesto that the critical failures in complex systems almost never happen because someone didn’t know what to do — they happen because the right question wasn’t asked at the right time. “Have you considered building an ACL?” is one of those questions. “What’s the exit strategy for the ACL?” is the other one. Ask both, up front, or you’ve only solved half the problem.

What Changes If You’ve Read the Book

There’s a quieter point underneath all of this, and it’s the one I keep coming back to.

Every person in that meeting with Somchai was experienced. Smart. Good at their jobs. None of them had reached for the ACL option before he raised it — including me. I’d lived the six-year migration. I knew viscerally why the standard playbook fails. But I hadn’t connected the experience to the vocabulary, which meant the tool wasn’t available when I needed it.

Somchai had read Domain-Driven Design. Not as a box to check — actually read it. And when he needed the frame, it was there.

This is not an argument that reading is better than experience. It’s an observation that they’re different paths to the same insight, and only one of them is transferable. The six-year migration made me a better engineer. It stayed with me. It will inform every system retirement conversation I have for the rest of my career. But it didn’t transfer — it didn’t help anyone in that room until I was able to tell the story. The vocabulary transfers. That’s the reason I write this post.

What to Actually Do

If you’re building a new system that’s replacing something with multiple clients:

Before the design is finalised, ask: how many clients does the old system have, what is their realistic migration timeline, and what is the weekly cost of supporting the old system while they migrate? Do the maths. Write it down. Make the trade-off explicit.

The equation is simple:

Cost of implementing the old contracts in the new system

vs.

Weekly cost of supporting the old system × number of weeks until the last client migrates

If the left side is smaller, build it. If the right side is smaller, hard cutover. The mistake isn’t choosing wrong — it’s never doing the calculation at all.

If an ACL is the right answer, treat it as first-class infrastructure from day one. Give it an owner — not the “person closest to it” but an actual named owner with accountability. Set a sunset date with teeth: a date after which the compatibility contracts will actually be removed, communicated far enough in advance that clients have time to move. If the sunset date passes without action, you haven’t built an ACL — you’ve built a second legacy system.

When you’re in the meeting and someone asks the question that makes the room go quiet — don’t reach for the whiteboard immediately. Sit with it. The discomfort you’re feeling is the plan being stress-tested, and that’s exactly what it needs.

The Bottom Line

Every system replacement project feels like it’s mostly a technical problem with a small coordination overhead. The ones that take six years felt that way too.

The migration tail is where plans go to die slowly. It’s not dramatic. Nobody calls a post-mortem on a migration that took four years longer than estimated — there’s no incident, no outage, just a quiet accumulation of ongoing cost that never made it onto the original estimate. The old system is still running. The engineer who knows how it works just accepted a job somewhere else. The deprecation announcement is three years old and nobody reads three-year-old announcements.

You don’t have to earn this lesson the hard way. The maths isn’t complicated. The pattern exists. The question — “what if the new system just spoke the old contracts?” — is available to anyone willing to ask it.

Ask it before you start.