Beer & Servers Don't Mix

The AI Coding Trap: Why 95% of Companies Are Measuring the Wrong Thing

Or: Your AI Can Write Code. So What?

There’s a statistic doing the rounds that should make every engineering leader pause: according to MIT’s 2025 report “The GenAI Divide,” 95% of companies investing in AI are seeing zero measurable return on their investment. Not poor returns. Not disappointing returns. Zero.

We’re collectively pouring $35–40 billion into AI initiatives while obsessing over the wrong metrics entirely.

As the economist Thomas Sowell once observed, “It is hard to imagine a more stupid or more dangerous way of making decisions than by putting those decisions in the hands of people who pay no price for being wrong.” We’ve handed the measurement of AI success to people counting lines of code while the actual value slips through their fingers like water through a sieve.

The Lines of Code Delusion

Here’s what I see in engineering teams everywhere, including my own: everyone’s excited about AI-assisted coding. Cursor, Claude Code, Copilot — pick your poison. Engineers are generating code faster than ever, and leadership is nodding along approvingly at productivity metrics that mean precisely nothing.

Let me be uncomfortably direct: writing code is maybe 10–20% of what engineers actually do. If you’ve optimised that slice and declared victory, you’ve essentially celebrated installing a more efficient engine in a car that’s still stuck in traffic.

Look at your actual workflow. It doesn’t start with typing code, does it? It starts with design discussions, feasibility assessments, architecture decisions. When your AI coding assistant fires up, the expensive thinking has already happened — or worse, hasn’t happened at all and now you’re generating code for the wrong solution faster than ever before.

And it doesn’t end with the commit either. Your CI pipeline fails. Someone needs to analyse the error. A form needs to be submitted to the translation system. An experiment needs to be configured in your A/B testing platform. A PR needs reviewing. Documentation needs updating. Stakeholders need notifying.

All of that work sitting on either side of code generation? That’s where the actual time goes. That’s where the actual value leaks out.

The Model Context Protocol Nobody’s Talking About

This is where MCP — Model Context Protocol — becomes genuinely transformative, and it’s baffling how few teams have noticed.

MCP lets AI assistants connect directly to your existing tools and systems. Not as some theoretical integration pattern, but as practical, working connections that turn your AI from a code-generating parlour trick into something that actually understands your workflow.

Here at Agoda, I’ve watched my teams transform their processes by thinking beyond the editor. The CI pipeline fails? With the GitLab or GitHub MCP, your AI doesn’t just write code — it pushes the code, monitors the pipeline, analyses the failure, and either fixes the issue or tells you why it can’t. One conversation, not three context switches.

We have engineers who used to manually submit forms to our translation system every time they added user-facing strings. Copy, paste, fill fields, submit, wait, come back later. Now? The AI that wrote the code also submits the translation request. Same with our A/B testing platform — creating an experiment used to mean leaving your IDE, logging into another system, filling out configuration forms, and hoping you didn’t fat-finger the targeting rules. Now it’s part of the same flow that wrote the feature.

The most useful one I saw was using the gitlab MCP and turned into a slash command in cursor /fix-mr-comments <Url to MR>. put this into cursor and it reads review comments, gets the right branch, addresses them in the code, comments back to the MR and resolves the comments as closed, the prompt is also instructed if there are questions in comments not changes to answer them, and if a suggestion is bad to push back against the reviewer.

This isn’t science fiction. This is MCP servers wrapping your existing APIs so your AI assistant can execute the manual tasks that were never really about the code in the first place.

Skills: The Missing Piece

The challenge with making AI genuinely useful across your tooling ecosystem is knowing which tools matter for which tasks. You don’t want your AI drowning in context about your translation system when you’re debugging a performance regression. This is where skills become essential — essentially curated sets of capabilities and context that you include only when they’re relevant to the job at hand.

Think of it as task-specific loadouts. Doing localisation work? Load the translation skill with its MCP connections and relevant prompts. Debugging CI? Load the pipeline skill. Running an experiment? Load the experimentation skill.

The teams getting actual value from AI aren’t the ones with the most tools connected — they’re the ones who’ve thought carefully about which tools matter for which workflows.

Beyond Generation: AI Code Review That Actually Works

We’ve gone further. We’ve put AI code review tools directly into our pipeline. The AI doesn’t just write the code; it reviews code too. Under certain conditions — and this is crucial, under certain conditions — we allow PRs that have been both written and reviewed by AI to merge without human intervention.

I can hear the sharp intake of breath from here. But consider: for well-defined, low-risk changes where the test suite is comprehensive and the AI reviewer understands the codebase patterns, why exactly do we need a human in the loop? The ceremony isn’t adding value; it’s adding delay.

The basketball coach John Wooden put it well: “Never mistake activity for achievement.” Having a human click ‘approve’ on every PR feels like work. Ensuring that every PR — regardless of who or what generated it — meets your quality bar before merging? That’s actually achievement.

Why 95% Are Getting Nothing

Let’s return to that brutal MIT statistic. Why are the vast majority of companies seeing zero return on their AI investments?

Because they’re measuring lines of code generated while ignoring the actual value stream.

Value doesn’t come from writing code faster. Value comes from shipping features that work, to customers who want them, faster than your competition. The code is an intermediate artifact. Nobody outside your engineering org cares about the code. They care about what the code enables.

When you expand AI’s scope from “generate code” to “participate in the workflow,” suddenly the value equation changes dramatically. That form submission that used to take five minutes? Gone. That context switch to check the pipeline? Gone. That manual experiment configuration? Gone. The cognitive load of remembering fifteen different systems? Dramatically reduced.

The 5% of companies seeing actual returns aren’t the ones with the best code generation metrics. They’re the ones who’ve reimagined their workflows around AI as a participant, not just a typing assistant.

The Bottom Line

We need to stop counting lines of code like they’re a meaningful measure of anything. As the physicist Richard Feynman once remarked about cargo cult science, we’re going through all the motions of AI adoption — the tools, the metrics, the dashboards — while missing the actual substance entirely.

The question isn’t “how much code is AI writing for us?” The question is “how much faster are we delivering value to customers, and how much friction have we removed from our engineers’ lives?”

MCP and agentic workflows aren’t just the next shiny thing. They’re the difference between AI as a toy and AI as a genuine force multiplier. The companies figuring this out will be the 5% seeing actual returns. The rest will keep celebrating productivity metrics while their competitors ship faster.

Your AI can write code. Congratulations. So can everyone else’s.

Now make it do the other 80% of the work.