The Migration That Never Ships: Why Your Team’s Best Intentions Are Paving the Road to Technical…
Or: Why 80% Complete Is the Most Dangerous Number in Engineering
The Migration That Never Ships: Why Your Team’s Best Intentions Are Paving the Road to Technical Hell
Or: Why 80% Complete Is the Most Dangerous Number in Engineering

It was 2:15 on a Thursday afternoon, and the senior tech lead sitting across from me hadn’t touched his coffee. He’d been at Agoda for years — survived multiple reorgs, shipped systems that handled millions of transactions, earned his stripes the hard way. I’d just finished outlining our next big migration plan, complete with timelines and team allocations. He leaned back, looked me dead in the eyes, and said: “Large migrations like that — they don’t work here.”
He wasn’t being dramatic. He wasn’t being difficult. He was being honest. And that honesty, uncomfortable as it was, turned out to be one of the most valuable things anyone has said to me in a planning meeting. Because when I dug into the history, he was right.
The 80% Trap
Seth Godin once said, “The only thing worse than starting something and failing is starting something and stopping.” He was talking about creative projects, but he could have been describing every major migration I’ve ever witnessed in a large engineering organisation.
Here’s the pattern, and if you’ve been in this industry long enough, you’ll recognise it with a sinking familiarity. A migration kicks off. Fifteen teams are involved, maybe more. Everyone’s bought in. There are biweekly updates for leadership, dashboards with progress bars, the works. The initial data looks phenomenal — performance gains, velocity improvements, engineers genuinely excited about the new stack. A quarter or two passes. Maybe three, depending on the size. High-fives all around. The metrics are real. The wins are real.
And then it stops.
Not dramatically. Nobody sends an email saying “we’re abandoning the migration.” It just… decelerates. The standup updates get vaguer. The tracking dashboard stops being updated. Other priorities emerge. And you’re left at 80% complete, which sounds close to done but is actually the worst possible place to be.
Because here’s what happened: you tackled the frequently changing, high-traffic areas first. Of course you did — that’s where the biggest wins were. The parts of the system that engineers touch every day, the hot paths that affect performance metrics leadership cares about. Those got migrated quickly and successfully. But the less frequently changed areas, the quiet corners of the codebase that nobody thinks about until they break — those got left behind. And if the architecture allows them to slip through, they will.
When You Can’t Slip Through, and When You Can
Not all migrations are created equal, and the architecture itself determines whether the long tail kills you.
Some migrations are all-or-nothing by nature. Major React version upgrades in a monolith, for instance — you can’t easily run multiple React versions in the same application. It’s a forcing function. You either finish or you don’t start. These migrations, paradoxically, tend to actually get completed because there’s no comfortable middle ground to settle into.
But when you can run things in parallel? That’s where the long tail grows. I’ve seen Objective-C to Swift migrations where both languages coexisted happily — until “happily” meant the Swift migration stalled at 75% because the remaining Objective-C code was in modules nobody wanted to touch. It worked, technically. It just never finished.
The same pattern plays out with framework and pattern migrations inside codebases. A new engineering leader arrives, decides on a better approach to handling DTOs, the team starts migrating. Then that leader moves on, a new direction emerges, and now you’ve got three different patterns for the same thing coexisting in the same codebase like geological strata — each layer a fossil record of someone’s best intentions.
The philosopher George Santayana warned us that “those who cannot remember the past are condemned to repeat it.” In our world, we don’t even need to forget — we just need to leave the past running in production alongside the present, and eventually nobody can tell which is which.
What Actually Works
I’m not going to pretend I’ve solved this. But I’ve learned enough from watching migrations succeed and fail to know what shifts the odds.
Measure progress and impact separately — but never lose track of progress.
This sounds obvious, but the failure mode is subtle. Teams naturally gravitate toward measuring impact because that’s what justifies the migration in the first place. Performance improvements, velocity gains, reduced error rates — these are the metrics that make leadership nod approvingly in quarterly reviews. But impact and progress are not the same thing, and conflating them is how you end up celebrating at 60% complete while the remaining 40% quietly rots.
For legacy system deprecation, look at traffic: old versus new.
This is better than any measurement in the code itself. You might think you’ve migrated a page, marked it done in your tracking spreadsheet, updated the Jira ticket. But users have bookmarks. There’s a deep link to the old version buried in support documentation that nobody’s updated in two years. For backend systems, you may not even know all your clients. An internal service you thought only had three consumers might have seven, because someone on another team started calling it directly eighteen months ago and never told anyone.
Traffic doesn’t lie. If requests are still hitting the old system, the migration isn’t done — regardless of what your project tracker says. Measure piece by piece, old versus new, and don’t close the book until the old traffic hits zero.
But traffic is a measure of progress, not impact. And this distinction matters for prioritisation. Your lower-traffic areas might actually be more valuable to tackle first depending on their impact on performance and velocity. The final steps in a conversion funnel, for example, are generally the lowest traffic but the most likely to impact conversion — users at that stage have higher intent. So use impact to drive priority at the start of a migration, and use progress to drive priority at the end, when the remaining work becomes housekeeping.
For pattern migrations in code, use static analysis and deprecation attributes.
When you’re migrating away from a coding pattern rather than a system, the measurement challenge is different. You need to track how many instances of the old pattern still exist and whether that number is going down over time.
Create static code analysis rules. Use deprecation attributes and annotations — [Obsolete] in C#, @Deprecated in Java and Kotlin — whatever your language provides. We have a CI template in GitLab that can consume output from several different linters, and we push it all into our data lake on merge to main. This gives us incremental change over time: a graph that should be trending toward zero. Tools like SonarQube can help here too, though custom rules might need some creative configuration.
Sometimes even raw filesystem metrics work. Counting file extensions in the codebase for language migrations is crude but effective — we did exactly this tracking a VB to C# migration years ago. Not sophisticated, but it told us what we needed to know.
Again, these are progress metrics. Don’t let them drive priority at the start — that’s impact’s job. Let them drive priority at the end, when the remaining work is completing what you’ve already started.
Smaller systems change the equation entirely.
Here’s where architecture becomes strategy. Smaller, more independent systems allow you to tackle migrations incrementally in a fundamentally different way than monoliths. A team can take a single system, migrate it completely, finish it, and measure all the benefits holistically. No long tail. No parallel running. Done.
This then creates something incredibly valuable: a working, proven example with real data. When the next team is deciding whether to invest in the same migration for their system, they’re not relying on benchmarks from the framework’s marketing page or some blogger running performance tests on a TODO list application. They’ve got a colleague’s actual production data showing what changed and by how much. That’s the kind of evidence that helps engineers make genuine cases for change — and also lets them know when the change isn’t worth the investment. Maybe for their particular system, the gains don’t justify the effort. That’s a valid outcome too, and proof from real data makes it a defensible decision rather than a guess.
The downside, of course, is inconsistency between systems. When each team migrates independently, you end up with different systems on different versions, different patterns, different approaches. But don’t fool yourself into thinking inconsistency is a solved problem in a monolith — it’s still there. The difference is visibility. In a monolith, inconsistency is staring you in the face every time you open the codebase. You don’t need special tooling to notice that three different DTO patterns coexist in the same repository. With smaller, independent systems, that inconsistency is spread across repos and teams, and you need deliberate tooling to even see it. Sourcegraph has been a godsend for us here — being able to search across every repository at once turns invisible drift back into something you can actually track and act on.
The Uncomfortable Truth
The tech lead who told me “large migrations don’t work here” wasn’t describing a unique pathology of our organisation. He was describing a universal pattern that most engineering leaders have lived through but few talk about openly, because admitting that your last three migrations stalled at 80% doesn’t make for a great conference talk.
The Scottish theologian William Barclay wrote, “Endurance is not just the ability to bear a hard thing, but to turn it into glory.” He wasn’t talking about software migrations, but he might as well have been. The teams that actually finish migrations aren’t the ones with the best kickoff decks or the most ambitious timelines. They’re the ones that treat the unglamorous final 20% — when the dashboards have stopped being interesting, when leadership has moved on to the next initiative, when the only people who care are the engineers still hitting the old code paths at 11 PM on a Wednesday — as the work that actually matters.
Start by measuring what matters: traffic for system migrations, static analysis for pattern migrations. Separate progress from impact so you know when you’re celebrating prematurely. Use impact to drive early priorities and progress to drive the long tail home. Build smaller systems where you can, so migrations have natural boundaries and produce real evidence for the teams that follow.
And if someone senior tells you that large migrations don’t work at your company? Don’t dismiss them. Buy them a coffee and ask them why. The answer will save you more time than any migration plan ever will.
Now, if you’ll excuse me, I need to go check the traffic dashboard on a system we “finished” migrating last quarter. I have a feeling the old endpoint is still getting hits from a support page nobody remembers writing.