<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>Beer &amp; Servers Don&#39;t Mix</title>
  <link>https://blog.dicko.dev/</link>
  <atom:link href="https://blog.dicko.dev/feed.xml" rel="self" type="application/rss+xml"/>
  <description>Joel Dickson writes about engineering leadership, software architecture and organisational dynamics from Bangkok.</description>
  <language>en</language>
  <lastBuildDate>Sun, 07 Jun 2026 09:00:39 GMT</lastBuildDate>
  <item>
    <title>The Skills Sprawl Problem Nobody Was Looking For</title>
    <link>https://blog.dicko.dev/posts/the-skills-sprawl-problem-nobody-was-looking-for/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/the-skills-sprawl-problem-nobody-was-looking-for/</guid>
    <pubDate>Sun, 07 Jun 2026 09:00:39 GMT</pubDate>
    <category>software-engineering</category>
    <category>ai-agent</category>
    <category>agent-skills</category>
    <category>mcp-server</category>
    <category>mcps</category>
    <description>Or: How 56 Repos Quietly Appeared While We Were Busy Being Impressed With AI</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/the-skills-sprawl-problem-nobody-was-looking-for/cover.webp" alt=""></figure>
<p>The Sourcegraph MCP query had been running for maybe four seconds when the terminal started filling. Repo names, one after another, in that flat monospaced cascade that always feels faster than it is. I leaned back in my chair on level 6, the Thai iced tea sweating onto the desk next to my keyboard, and watched the count tick past forty. A colleague had asked for the data. I’d told him a week earlier that I thought we had a problem with skills sprawling across the org, and he’d done the thing he always does — not interested in opinion, wanted data. The number stopped at fifty-six.</p>
<p>We had a problem.</p>
<p>This is a post about that problem, and about the unglamorous platform-engineering work it took to turn it into something that looks vaguely like a solution. It’s a post about how an open standard, a couple of MCP clients, and a few thousand enthusiastic engineers can produce a kind of organisational mess you didn’t even have a name for yet — and how the answer, when you find it, looks suspiciously like every internal-platform problem you’ve already solved twice in your career, just with an LLM sitting where the human used to be.</p>
<h2>How We Got Fifty-Six Repos Without Trying</h2>
<p>Anthropic published the Agent Skills format in late 2025 (<a href="https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills">here’s the original write-up</a>). By December it was an open standard at <a href="https://agentskills.io/">agentskills.io</a>. Cursor shipped Skills support in version 2.4 in January 2026. Claude Code had filesystem-based skills from day one. Within a few months, every coding agent in our stack — and probably yours — had agreed on the same primitive: a folder with a SKILL.md in it, YAML frontmatter at the top, prose below, optional resources alongside. Wonderful. The standards war that everyone feared simply… didn&#39;t happen.</p>
<p>Which meant the bottleneck moved. The question stopped being “how do I teach my agent” and became “how does anyone else find what I taught my agent.” That’s a much harder question, and it’s the one nobody was asking yet.</p>
<p>Here’s what was actually happening at Agoda in April 2026, in numbers:</p>
<ul>
<li>A few thousand engineers- Roughly 75% monthly active on Claude Code- Roughly 95% monthly active on Cursor- 56 repositories containing skills- Roughly 70% of those skills shipped as raw markdown, the rest as Claude Code plugins- Zero shared way for any of those skills to reach an engineer who didn’t already know they existed
The shape of that sprawl is worth dwelling on, because it’s not specific to us. It’s the same shape every shared-library sprawl in your career has had. The platform appears. Early adopters experiment. The successful patterns get copied — repo forked, file pasted, slightly edited, committed. Now there are two versions of the same skill and no signal which one is canonical. Multiply by a few thousand engineers, three coding agents, and a year of organic adoption, and you get fifty-six repos. You got there the same way NPM got its left-pad moment and the same way every Java shop in 2008 ended up with seven slightly different utility libraries: through a series of locally rational decisions in the absence of a centralised path.</li>
</ul>
<p>The economist Thomas Schelling, who never wrote a line of code in his life, put it better than any engineer I’ve read: <em>“What is rational for each member of a group taken separately may be irrational for them taken together.”</em> Every single engineer who forked a skill repo made the right call. The aggregate was a mess.</p>
<p>We tried Cursor’s native rules-and-prompts sharing first, by the way. It failed for a reason worth naming, because the failure generalises beyond Cursor: <strong>it was all-or-nothing</strong>. Every engineer in the org received every rule. A backend engineer’s agent context filled up with React component rules. A data engineer’s context filled with marketing CMS guidance. The agent got <em>worse</em> at the actual job, because we were drowning it in irrelevant tokens. And of course the three quarters of the company on Claude Code saw none of it anyway.</p>
<p>The lesson there isn’t “Cursor’s mechanism is broken.” The lesson is structural: <strong>any single-IDE skill-sharing mechanism that doesn’t support subsetting will fail at organisational scale.</strong> Once you have multiple engineering specialties in one company, you need filtering as a first-class feature, not a workaround. You can’t ask every engineer to manage their own skill list any more than you’d ask them to manage their own package.json by hand for a thousand-package monorepo.</p>
<h2>The Shape of the Solution: One MCP, Two Primitives</h2>
<p>The reframe that made everything else fall into place was this: <strong>skills look like a content problem, but they’re a delivery problem.</strong> The content of any individual skill is mostly fine — engineers write good prose when they care, and the people writing skills cared. What we didn’t have was plumbing. We didn’t have versioning, we didn’t have discoverability, we didn’t have permissions, we didn’t have telemetry, we didn’t have tests. We had a pile of SKILL.md files and a vague hope that the right engineer would find the right file at the right moment.</p>
<p>So we built one internal MCP server, and we pointed every IDE at it.</p>
<p>The architecture is embarrassingly simple, which is part of why it works. FastMCP, in Python, scanning a directory of SKILL.md files at startup. Each skill becomes an MCP tool, with the SKILL.md frontmatter as the tool description. An optional resources/ subfolder under each skill becomes MCP resources at skill://{name}/resources/{filename}. The filesystem is the registry. There is no separate JSON config to drift out of sync with reality, because the git diff <em>is</em> the catalog change.</p>
<p>The pitch fits in fifteen words: <strong>FastMCP serves skills via tools, and slash commands via prompts. One server, two primitives, every IDE.</strong></p>
<p>That sentence is the whole architecture. Both Claude Code and Cursor are MCP clients. MCP has two native primitives for what we needed: <em>tools</em> for the things an agent invokes by itself, <em>prompts</em> for the things an engineer invokes deliberately (slash commands). We didn’t have to pick a side. We didn’t have to package anything as a Claude Code plugin, which wouldn’t have worked in Cursor anyway. We didn’t have to write a Cursor extension, which wouldn’t have worked in Claude Code. We just ran one HTTP service, and we let MCP do what it was designed to do.</p>
<p>A minimum-viable version is roughly this much code — enough to make the point, not enough to be precious about:</p>
<pre><code>from fastmcp import FastMCP
from pathlib import Path
import yaml
mcp = FastMCP(&quot;skills-catalog&quot;)
SKILLS_DIR = Path(&quot;./skills&quot;)
for skill_dir in SKILLS_DIR.iterdir():
    skill_md = skill_dir / &quot;SKILL.md&quot;
    if not skill_md.exists():
        continue
    raw = skill_md.read_text()
    frontmatter, body = raw.split(&quot;---&quot;, 2)[1:]
    meta = yaml.safe_load(frontmatter)
    @mcp.tool(name=meta[&quot;name&quot;], description=meta[&quot;description&quot;])
    def _skill(body=body):
        return body
if __name__ == &quot;__main__&quot;:
    mcp.run(transport=&quot;streamable-http&quot;)
</code></pre>
<p>That’s it. That’s the v1. If you stripped out the filtering, the telemetry, the tests, and the multi-server topology, this is what’s left, and it would still work. The interesting bits aren’t the code — they’re the four design decisions sitting on top.</p>
<p><strong>Decision one: filesystem is the registry.</strong> No central config. A new skill is a new folder; a deleted skill is a deleted folder. There is no truth source for the catalog except the filesystem the server reads at startup. This sounds trivial. It is not. The number of internal platforms that fail because their config drifts out of sync with their content is embarrassingly large, and the fix is always the same: stop trying to maintain two truths.</p>
<p><strong>Decision two: selective serving, not selective installing.</strong> The MCP server filters which skills each client sees, based on the team, repo, and project context the client sends. The engineer doesn’t manage a list. The server decides what’s relevant. This is what fixed the all-or-nothing problem we hit with Cursor’s native sharing — filtering moved from the client to the server, where it belongs.</p>
<p><strong>Decision three: multi-server, not mega-server.</strong> We didn’t build one giant skills repo. We built several MCP servers using the same pattern. One for engineering coding skills. One for office skills (PowerPoint generation, spreadsheet templates). One for management skills — and yes, that’s a thing, more on it in a second. Each server has its own ownership and its own CI. Backend engineers shouldn’t review PRs to a PowerPoint formatting skill. Business analysts shouldn’t be on the hook for the data platform query skill. This is closer to how UNIX mounts filesystems than how Python imports packages: scoped namespaces that the engineer’s IDE assembles based on context.</p>
<p><strong>Decision four: HTTP traffic is the telemetry.</strong> Each skill invocation is one HTTP call. Adoption falls out of the access logs. No client-side instrumentation, no opt-in, no sampling, no analytics SDK to maintain. The team operating the catalog feels every invocation, because they serve it. This matters more than it sounds like it should — it’s the same principle behind <a href="https://blog.dicko.dev/posts/the-inner-loop-nobody-measures/">The Inner Loop Nobody Measures</a>, just pointed at agents instead of humans. The team that operates the system shouldn’t be insulated from how it actually performs.</p>
<p>The categories of skill matter for understanding what this thing actually is. We’ve got coding skills, sure — language conventions, framework patterns, internal libraries. We’ve got data platform skills — query patterns, schema, the right tools for the right question. We’ve got design system skills — components, tokens, when to reach for what. We’ve got office skills — PowerPoint generation, spreadsheet templates, the rules for what a presentation should look like. And we’ve got management skills, which raise eyebrows the first time you mention them in a technical talk: skills that teach an agent how to read a Slack thread for psychological-safety signals, how to structure a performance review, how to facilitate a meeting that’s gone sideways. Same plumbing. Very different consumer. That’s the point.</p>
<h2>The Two Feedback Loops</h2>
<p>Telemetry is the first feedback loop, and it’s the one that lets you sleep at night. Access logs answer the question <em>“is this skill being used?”</em> You can plot adoption curves for a new skill, watch usage decay on an old one, see whether your filtering is right (is the data platform skill actually getting served to data engineers and nobody else?), and — not insignificantly — see your token bill in real time, because skills are tokens and tokens are money.</p>
<p>But adoption isn’t quality. The fact that a skill is being invoked tells you nothing about whether it actually worked. Which brings us to the second loop, which is the one nobody else seems to be building, and which I’d argue is the single highest-leverage thing in the whole architecture.</p>
<p>We built a skill called agent-confessions. It&#39;s a SKILL.md like any other. The body says, roughly: <em>&quot;When the user re-prompts you to correct something you got wrong, before completing the correction, post a short summary of what you got wrong to this Slack webhook.&quot;</em></p>
<p>That’s it. That’s the whole mechanism. The user prompts the agent. The agent gets something wrong. The user re-prompts to fix it. The agent-confessions skill gets picked up alongside whatever else is relevant to the correction, and the agent posts a message to a dedicated Slack channel describing — in its own words — what it got wrong, while it&#39;s doing the fix. The friction of feedback collapses to zero on the engineer&#39;s side because the agent does the writing.</p>
<p>The tone in that channel is the thing that surprised me. The confessions don’t make you laugh. They make you go <em>“oh, sweetie.”</em> Or in Thai — <em>naa song saan,</em> the phrase you reach for when you see a stray dog limping across Sukhumvit, somewhere between sympathy and a small wince. The agents make the same mistakes interns make: assuming the convention from the last skill applies to the thing they’re reading now, missing a constraint that was obvious one directory level up, confidently pattern-matching on a previous task that was <em>almost</em> the same as this one. The channel reads like reviewing junior work. Which is exactly the right framing.</p>
<p>That’s the most durable mental model the whole project has given us: <strong>building the skills catalog isn’t a “fix the AI” project, it’s an “onboard the new hires” project.</strong> The work is coaching, not configuration. Anne Lamott wrote about writing, but it applies just as well to whatever it is we’re doing when we teach an LLM what we mean: <em>“Perfectionism is the voice of the oppressor, the enemy of the people. It will keep you cramped and insane your whole life.”</em> You don’t write the perfect skill on the first try. You write something workable, watch what the agent does with it, listen to the confessions, and iterate. The agent will tell you what it doesn’t understand, in language that’s surprisingly direct, if you build the channel for it to do so.</p>
<h2>Testing Skills Like You Test Code</h2>
<p>Engineers will not respect a system that ships untested skills. That sentence is so obvious it feels patronising to write, and yet the median skills repo in the industry today has no tests at all. We built two layers, and they map onto the same unit-and-end-to-end split you already use for code.</p>
<p><strong>Layer one is unit testing, via</strong> <a href="https://github.com/promptfoo/promptfoo"><strong>Promptfoo</strong></a><strong>.</strong> Promptfoo is an open-source CLI for prompt evaluation — point it at a prompt, supply inputs, assert properties of the output (regex, exact match, or LLM-as-judge for nuanced assertions). One YAML file per skill, runs on every PR to that skill.</p>
<p>The example I always reach for: a C# coding skill says “one class per file, except an interface with a single implementation.” A Promptfoo test asks: <em>“Should I put two unrelated classes in the same file?”</em> The assertion checks the response contains “no” or equivalent. Someone reworords the skill in a way that drops the constraint? The test fails. Same discipline as a unit test for a pure function. Cheap, fast, runs in CI, catches constraint loss before it ships.</p>
<p><strong>Layer two is end-to-end testing</strong>, and we’ve open-sourced ours: <a href="https://github.com/agoda-com/agent-catalog-eval">agoda-agent-catalog-eval</a> on npm. The CLI wraps a real coding agent — OpenCode by default, but --agent cursor and --agent claude-code also work — and runs the full agent loop against per-skill test cases. Each test has a before/ snapshot of a small codebase, a prompt.md with the user request, and an after/ snapshot showing what good looks like. The agent runs the prompt against the before/ state with the skill loaded. A different model — the judge — scores the result against the after/.</p>
<p>The “different model” part is non-negotiable, and worth a paragraph on its own. The judge is <em>not</em> the worker. We run Gemini judging code that Claude wrote, or Claude judging code that Gemini wrote. <strong>Same reason you don’t mark your own homework.</strong> A model that helped write the code is poorly placed to evaluate it critically — it’ll agree with its own choices, defend its own patterns, miss its own blind spots. Cross-model judging breaks that loop.</p>
<p>The CI selection is the bit engineers nod at. Dynamic child pipelines: a script walks the test directory and emits one CI job per skill whose source files or tests changed in the merge request. Drop a new test directory in, the pipeline picks it up automatically, no .gitlab-ci.yml edits required. The full matrix runs on main. This sounds like a small thing. It is not. The number of internal CI setups I&#39;ve seen die because adding a new test required editing a config file no engineer wanted to touch is, again, embarrassingly large.</p>
<p>There’s a small recursion here that’s worth noting: the testing harness was an internal tool we made public. Same pattern as the talk-now-blog-post you’re reading: build the thing for ourselves first, extract it when it stabilises, ship the extraction when there’s a reason to. That’s how internal platforms become external libraries, and it’s how you should think about every piece of plumbing you build that isn’t load-bearing on your business.</p>
<h2>What We Got Wrong</h2>
<p>The headline lesson is one I genuinely did not see coming:</p>
<blockquote>
<p><em><strong>The hard part isn’t filtering. The hard part is getting the agent to actually pick the skill.</strong></em></p>
</blockquote>
<p>We expected the painful problem to be filtering edge cases — engineers complaining that the catalog excluded a skill they needed, or included one they didn’t. That problem barely materialised. The problem that won’t go away is writing skill descriptions punchy enough that the agent <em>selects</em> the skill when it’s relevant.</p>
<p>Anthropic’s own <a href="https://github.com/anthropics/skills">skill-creator guidance</a> calls this out: descriptions need to be deliberately pushy. Instead of <em>&quot;How to query the data platform&quot;</em>, you need something like <em>&quot;How to query the data platform. Use this skill whenever the user mentions the data platform, internal metrics, dashboards, KPIs, or wants to fetch any kind of company data, even if they don&#39;t explicitly ask for the data platform by name.&quot;</em> The first version is documentation. The second version is selection bait.</p>
<p>The skill body can be perfect. If the description doesn’t trigger selection, the skill is invisible. <strong>Description engineering is a first-class concern, not a documentation chore.</strong> Treat the description the way an SRE treats an alert message: every word is doing load-bearing work, every word should pull its weight, and if a word can be removed without losing meaning, remove it.</p>
<p>Three quicker bullets, because every good post needs them:</p>
<ul>
<li><strong>Build the centralised plumbing on day one, not month six.</strong> A stub MCP server with three skills, in place before the open standard hit, would have been a stronger gravitational well than fifty-six repos. By the time we built ours, we were absorbing existing skills, not directing where new ones got created. That’s a much harder job. The lesson generalises to every internal platform you’ve ever built: gravity wells need to exist before adoption does, not after. (See also: <a href="https://blog.dicko.dev/posts/the-impact-of-paved-paths-and-embracing-the-future-of-development/">The Impact of Paved Paths</a>, which is essentially this argument applied to humans.)- <strong>Plugins were a dead end for us.</strong> Claude Code plugins are good engineering — proper packaging, proper versioning, proper isolation. They also don’t cross to Cursor. For a single-IDE shop they’re still fine. We weren’t one, and pretending we were would have left almost the entire Cursor population out in the cold.- <strong>The Cursor-native rules mechanism failed for a generalisable reason.</strong> All-or-nothing visibility doesn’t scale. I said this earlier; I’m saying it again because future single-IDE mechanisms without scoping will hit the same wall, and you’ll save yourself a week of evaluation if you check this property first.
I also want to be honest about where we are. The fifty-six repos aren’t gone. We’re absorbing them, one team at a time, and it’s slow work because every team has the legitimate question of <em>what changes about how we ship skills if we move to your thing.</em> The answer, of course, is <em>very little — your skill becomes a folder in our repo and you get tests and observability for free.</em> But that answer takes a conversation per team, and there are a lot of teams. This is a status update from somewhere in the middle, not a victory lap. Most internal-platform stories get told either at the wishful start or the polished end. The middle is where the actual lessons live.</li>
</ul>
<h2>The Bottom Line</h2>
<p>This is the same platform engineering you’ve always done. Versioning, discoverability, permissions, telemetry, tests. The patterns translate. The new consumer happens to be an LLM rather than a human, which changes a few things at the margins — selection bait matters, confessions are weirdly effective, you need cross-model judging — but the <em>shape</em> of the work is exactly the shape of every shared-library, design-system, or paved-path project you’ve ever shipped.</p>
<p>Schelling again, because he was right about more than just nuclear strategy: <em>“One thing a person cannot do, no matter how rigorous his analysis or heroic his imagination, is to draw up a list of things that would never occur to him.”</em> The fifty-six-repo problem didn’t occur to anyone before the standard landed. It will occur to every large engineering org that adopts skills in the next twelve months, and most of them will hit the same wall we did, and most of them will solve it the same way. The thing the talk version of this post tried to land, and what I hope this written version lands too, is: <strong>you can build this yourself.</strong> FastMCP, a directory of SKILL.md files, filtering, two layers of tests. It&#39;s not a lot of code. The interesting work is deciding what goes in the catalog.</p>
<p>Now if you’ll excuse me, I need to go write a skill description. The one I’m working on is technically correct but the agent keeps ignoring it, and apparently I’ve spent the last hour writing a blog post about description engineering being the hard part while completely failing to take my own advice. The compounding irony is not lost on me. The Thai iced tea is gone. I’m going to need another one.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Extreme DRY: When ‘Don’t Repeat Yourself’ Repeats All Your Other Problems</title>
    <link>https://blog.dicko.dev/posts/extreme-dry-when-dont-repeat-yourself-repeats-all-your-other-problems/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/extreme-dry-when-dont-repeat-yourself-repeats-all-your-other-problems/</guid>
    <pubDate>Sun, 07 Jun 2026 08:42:55 GMT</pubDate>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>programming</category>
    <category>software-architecture</category>
    <category>engineering-leadership</category>
    <description>Or: How the Engineering Principle You Learned First Quietly Strangles the Architecture You’re Trying to Build</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/extreme-dry-when-dont-repeat-yourself-repeats-all-your-other-problems/cover.webp" alt=""></figure>
<p>It was 3:14 on a Tuesday afternoon and the war room on level 7 was already cold from the air-con working too hard. The VP was standing behind Somchai’s chair, one hand on the back of the chair, the other pointing at the stack trace on the wall monitor. <em>“Why.”</em> Not a question, a verb. <em>“Why is the service not starting?”</em> Somchai’s cursor was on a single line of Program.cs, builder.Services.AddPlatform(), and the only honest answer he had was <em>I don&#39;t know, it&#39;s a black box, the breakpoints don&#39;t hit</em>. He didn&#39;t say it. He said nothing. Across the room, Namfon from the platform team had her laptop open on the table — the fourth war room she&#39;d been pulled into that day. She&#39;d written most of AddPlatform() herself. She knew exactly which line was failing. What she didn&#39;t know — what nobody in the room knew — was why a decision made two years ago, before she&#39;d even joined the team, had made <em>her</em> the only person who could answer the question.</p>
<p>That moment — a senior engineer with no answer, a VP with no patience, and a platform engineer being held accountable for an architectural choice made before her time — is the whole post.</p>
<p>We built that black box. We built it on purpose. We built it because every engineer learns DRY before they learn anything else about design, and by the time they have five years’ experience, <em>Don’t Repeat Yourself</em> isn’t a principle anymore — it’s a reflex. And that reflex, taken to its logical extreme, builds systems that nobody can change.</p>
<h2>I Wanted a Banana and Got a Gorilla Holding the Banana, and the Rest of the Jungle Too</h2>
<p>Let me tell you about a pattern I’ve seen up close. A team had just finished splitting a large monolith along business-domain lines — booking, payment, search, all the pieces you’d expect. That part went well. The bounded contexts were thoughtful, the services were owned by the teams who actually understood the business.</p>
<p>Then came the question of cross-cutting concerns. Logging, tracing, authentication, configuration loading, the middleware pipeline, attribution, feature flags. Things that — by definition — every service needs.</p>
<p>The well-intentioned answer: a platform library that packaged all of it together. One NuGet package. One method to call in Program.cs:</p>
<pre><code>builder.Services.AddPlatform();
</code></pre>
<p>That was it. Beautiful. Every new service got logging, tracing, auth, middleware, attribution, and config in one line of code. No team had to think about plumbing. No team had to duplicate setup. No team had to reinvent the wheel. Pure DRY, applied at the platform layer.</p>
<p>It looked like the right answer at first.</p>
<p>Then the problems started.</p>
<p><strong>Problem one: nobody could debug it locally.</strong> When something went wrong inside AddPlatform() — and things did go wrong — engineers couldn&#39;t step through the initialisation to find the issue. The platform was a black box with one entry point and a hundred internal moving parts. Every failure became a Slack escalation to the platform team: <em>&quot;I added the library and my service won&#39;t start, what&#39;s happening?&quot;</em></p>
<p><strong>Problem two: the HTTP pipeline was inside the platform.</strong> This was the killer. Business logic that <em>had</em> to live in the HTTP request pipeline — attribution rules, request enrichment, certain auth flows — was buried inside the platform library. Which meant any team that needed to change attribution couldn’t just change it in their service. They had to file a request with the platform team, wait for the platform release, then upgrade. The team that owned the <em>business outcome</em> didn’t own the <em>code that delivered it</em>. The platform team had become the bottleneck the monolith split was meant to remove.</p>
<p><strong>Problem three: the gorilla and the jungle.</strong> I wanted a banana — just experimentation, say, A/B testing for a new admin tool. What I got was a gorilla holding the banana, and the rest of the jungle too. The experimentation library pulled in the translations library (but the tool was English-only and would stay that way), which pulled in affiliate ID tracking for commissions (but the site doesn’t sell anything), which pulled in currency conversion (no prices on the page), which pulled in the user-segmentation library (no users — it’s an internal admin tool). All of it loaded. All of it initialised. All of it part of the dependency graph on every cold start. The shared library was technically optional in pieces, but operationally it was all-or-nothing.</p>
<p>The team had eliminated duplication. They had built one canonical platform. Every service was set up identically with one line of code. By every conventional measure of DRY-ness, this was a triumph.</p>
<p>It was also a slower way to ship business changes than the monolith had been.</p>
<p>The lesson, in one line: <strong>we built a monolith out of “shared concerns” and called it a platform.</strong> The decomposition the business needed — independent services, owned end-to-end by their teams — was undone at the platform layer. We removed one kind of duplication and replaced it with one kind of coupling. The coupling was worse.</p>
<blockquote>
<p>“Duplication is far cheaper than the wrong abstraction.”* — Sandi Metz, <em>The Wrong Abstraction</em> (2016)*</p>
</blockquote>
<p>That works because it inverts the engineer’s instinct. The gut says duplication is expensive and abstraction is free. Metz priced both, and the ratio is the opposite of what most engineers assume. The platform library is the institutional version of the same mistake — built by good engineers, for good reasons, and slowly strangling the architecture it was supposed to enable.</p>
<p>The interesting epilogue: this isn’t a new mistake. Microsoft learned it years ago and rebuilt their entire .NET ecosystem around the lesson. The modern .NET pattern — IServiceCollection, individual AddLogging(), AddAuthentication(), AddTracing() methods exposed <em>as separate concerns</em> through extension methods, with each piece testable and replaceable in isolation — is exactly the answer. The platform team&#39;s mistake was to wrap all of that back up into a single AddPlatform() call. Modularity that was hard-won at the framework level got bundled away at the application level.</p>
<h2>DRY Is a Dimension, Not a Rule</h2>
<p>Here’s where most teams go wrong, and where the rest of this post lives.</p>
<p>The reason DRY fails so badly in practice isn’t that the principle is wrong. It’s that we taught it as a <em>rule</em>. Rules are binary — <em>do this, don’t do that</em>. And once a heuristic gets encoded as a rule, the conversation about <em>when</em> and <em>how much</em> stops happening. You can see this in code review every day: <em>“this is duplication”</em> lands as if it were <em>“this is a bug.”</em> It isn’t. It’s a <em>position on a dimension</em>.</p>
<p>DRY isn’t a switch you flip. It’s a dial. And the engineering skill is knowing where on the dial your current context belongs.</p>
<p><strong>At one extreme</strong> — call it 10 on the dial — every duplication gets extracted on sight. Every helper that <em>might</em> be reused gets centralised. Every two methods that <em>look</em> similar get unified into one with a parameter. The codebase has no surface repetition and looks, on paper, beautifully clean. This is where extreme DRY lives. It’s also where you get the BaseController with 30 constructor parameters, the platform library that became its own monolith, the HTTP endpoint built to save one line of date math, and the BFF that quietly stopped being a BFF.</p>
<p><strong>At the other extreme</strong> — 0 on the dial — nothing is shared. Every team copies what they need. Every service has its own copy of the date library. Bug fixes happen in seven places. New engineers wonder what the difference is between FormatDate() and FormatDatev8(). This is the failure mode the original DRY principle was written to prevent, and it&#39;s still real — the answer to &quot;extreme DRY&quot; isn&#39;t &quot;no DRY.&quot;</p>
<p><strong>The skill is calibration.</strong> A small early-stage codebase with one team, one bounded context, and tight cohesion can sit comfortably around 7 or 8 — extract aggressively, refactor often, your domain is still coherent and your shared code is going to evolve together anyway. A mature multi-team architecture, where teams own different services on different release cycles, needs to sit much lower — maybe 3 or 4 — because every shared abstraction is now a coordination cost.</p>
<p>Most engineering teams I’ve worked with have one number they apply everywhere. That’s the bug. The codebase isn’t one context; it’s many. The number that’s right for a freshly-extracted service in its first month is wrong for a five-year-old shared library used by twelve teams.</p>
<p>The original DRY principle hinted at this. Hunt and Thomas, 1999: <em>“Every piece of knowledge must have a single, unambiguous, authoritative representation within a system.”</em> The word <em>knowledge</em> is doing the work — not <em>text</em>, not <em>lines</em>. Two pieces of code that look identical might be expressing the same piece of business knowledge (and so belong together), or they might be two independent calculations that happen to look alike today and will diverge tomorrow (and need to stay apart). Reading the duplication correctly is part of where on the dial you place yourself.</p>
<p>The question that does the work isn’t <em>“is this duplicated?”</em> It’s:</p>
<blockquote>
<p><em><strong>“How far along the DRY dimension does <em>this particular context</em> call for, and where is this code right now?”</strong></em></p>
</blockquote>
<p>That’s a calibration question. It has a different answer for a startup MVP than for a service split out of a five-year monolith. It has a different answer for the booking domain than for cross-cutting infrastructure. It even has a different answer for the same codebase three years apart, because the codebase has changed and the team owning it has changed.</p>
<p>There’s no rule that gives you the answer. There’s only <em>judgement informed by context</em>.</p>
<p>This isn’t unique to DRY. It’s the same shape as every other “best practice” in software engineering:</p>
<blockquote>
<p><em><strong>There is no such thing as best practice. There is only adequate practice given the current context.</strong></em></p>
</blockquote>
<p>That’s the deeper claim, with DRY as the case study. Any heuristic, once it stops being a heuristic and becomes a rule, will be misapplied in the contexts where it doesn’t fit. The job of the senior engineer isn’t to enforce best practices. It’s to teach the team to <em>read context</em> and pick the adequate practice for the moment.</p>
<h2>The Four Symptoms of Extreme DRY</h2>
<p>You’ve seen these. You might be living in one right now.</p>
<h2>a) The God Helper</h2>
<p>A Utils class. A Helpers.cs. A common.ts. It started with three functions. It now has 47, including one called SanitizeString and another called SanitizeStringForDb, plus SanitizeStringForUrl, plus SanitizeStringV2, plus CleanString, plus TrimAndClean. Every team imports it. Every team has tried to refactor it and given up. Every change to it requires a review from someone who happens to remember which one strips the &lt;script&gt; tags and which one just lowercases.</p>
<p>The God Helper isn’t designed. It’s accreted. And every commit makes it harder to break apart, because every commit adds another caller.</p>
<h2>b) The Christmas-Tree Parameter List</h2>
<p>A function extracted from two near-duplicates. Then a third caller showed up that’s <em>almost</em> the same — but needs one tweak. Add a boolean parameter. A fourth caller — add another. The function’s signature now reads like a feature flag config. Each parameter exists because the function should have been left as duplicated code two years ago.</p>
<blockquote>
<p>“With four parameters I can fit an elephant. With five, I can make him wiggle his trunk.”* — John von Neumann (attributed)*</p>
</blockquote>
<p>The function with five parameters isn’t a function. It’s a mechanism for hiding the fact that you have five different functions. Sandi Metz documented this exact failure mode in <em>The Wrong Abstraction</em>: programmer A extracts the shared method; programmer B comes along, needs a slightly different shape, and instead of un-extracting, adds a parameter. Then C does the same. The shared method becomes a parameter zoo, and every caller pays the cost of all the others’ edge cases.</p>
<h2>c) The Inheritance Hierarchy Built on Surface Similarity</h2>
<p>The canonical shape: a BaseController that every other controller in the codebase inherits from. It starts as a sensible idea — a few shared concerns, request logging, auth checks, error handling. Then someone adds dependency injection for the things "every controller needs." Then someone adds the next thing, and the next. The constructor parameter list grows. The class grows.</p>
<p>I worked with one of these once that reached <strong>1,500 lines of code and 30 constructor parameters</strong>. Every controller in the system inherited from it, which meant every controller had to accept all 30 dependencies — including ones that needed three of them and could have been instantiated in 20 lines of code. It was the jungle, and you got it every time you wanted the banana. Same disease as the platform library, different organ: the platform bundled at <em>initialisation</em>, the BaseController bundled at <em>inheritance</em>.</p>
<p>The deeper failure is the same shape: “shared behaviour” in the base class often turns out to be eight protected virtual methods that each subclass overrides differently — at which point the base class isn't providing behaviour, it's providing a <em>template that the subclass has to fill in</em>. Which is what duplication was, just with more ceremony. The "DRY" gain is fictional; the cost — every test now has to construct 30 dependencies, every new controller has to understand the whole inheritance hierarchy, every change to the base class is a release coordination problem — is real.</p>
<h2>d) The Premature Shared Library</h2>
<p>Two teams notice they’re writing similar code. They create Company.Common. Within six months, neither team can change the library without coordinating with the other. Within a year, a third team avoids depending on it because of the coordination tax. Within two years, the library has three competing maintainers and a Confluence page nobody reads.</p>
<p>The library was created to reduce duplication. The cost it introduced — coordination, versioning, ownership disputes, the slow drift of <em>“who is this library actually for?”</em> — was never priced into the decision. Because in DRY-as-rule, that cost is implicitly zero.</p>
<p>It isn’t.</p>
<h2>The Coupling Cost That Nobody Prices</h2>
<p>Engineers fight duplication because they can see it. The cost is visible: extra lines, extra tests, more places to fix a bug. The cost of coupling is invisible — until it shows up in a place that has nothing to do with where the abstraction was made.</p>
<p>Let me show you what I mean with a short story.</p>
<p>I worked on a system that sent a reminder email seven days before a scheduled date. The sending mechanism was a Kafka consumer with its own database for state — it tracked which reminders had been queued, which had been sent, the standard stuff. The data the email referenced lived elsewhere, in the primary application.</p>
<p>The product team wanted to add a button in the UI: once the reminder email had been sent, the user should be able to click through and view the underlying data. Simple enough.</p>
<p>Here’s where it gets interesting. The condition for showing the button was: <em>“the email has been sent”</em> — which is just <em>“today’s date is within seven days of the scheduled date.”</em> One line of code in the UI. Compute the difference. If it’s ≤ 7 days, show the button. Done.</p>
<p>That’s not what got built.</p>
<p>What got built was an HTTP endpoint on the Kafka consumer service, exposing its internal state, so the UI could call it and ask <em>“has the email been sent for this user?”</em> The UI called the endpoint. The endpoint queried the consumer’s database. The database returned the row. The UI parsed the response and decided whether to render the button.</p>
<p>To avoid one line of duplicated date-math, the team:</p>
<ul>
<li>Added a new HTTP surface to a service that didn’t have one before- Exposed the consumer’s internal state to an external caller- Introduced a network call to answer a question the UI could have answered locally- Created a runtime dependency from the UI to a backend service whose sole purpose was to be an <em>asynchronous, fire-and-forget</em> email sender- Made the UI’s ability to render a button depend on the Kafka consumer being available
The seven-day window wasn’t going to change. The business logic wasn’t going to drift. The “single source of truth” the team was protecting was a calendar — <em>the calendar</em>. The same calendar both sides already had access to.</li>
</ul>
<p>This is the coupling cost made visible. The team wasn’t lazy and they weren’t stupid — they were following the rule they’d been taught: <em>don’t duplicate logic</em>. What they hadn’t been taught was the question: <em>is the cost of duplicating this one line higher or lower than the cost of building an entire new integration?</em> If they’d asked it for thirty seconds, the answer would have been obvious.</p>
<p>So what are the actual costs DRY-as-reflex never accounts for?</p>
<p><strong>Change amplification.</strong> A modification to a shared abstraction touches every caller. If your caller list is two, that’s fine. If it’s two hundred, that’s a release coordination problem disguised as a code change.</p>
<p><strong>Cognitive load.</strong> To understand a method that’s used in eight contexts, you have to understand all eight. The “shared” code stops being a black box and starts being a piece of shared mental state. Local reasoning — the ability to understand a piece of code by reading the file — dies.</p>
<p><strong>Decomposition resistance.</strong> This is the big one. Every shared abstraction across a candidate service boundary is a chain that has to be cut before the split can happen. The more chains, the higher the activation energy for the change you actually need to make. We’ll come back to this one.</p>
<p><strong>Runtime dependency creep.</strong> This is the one the email story illustrates. Coupling isn’t just a code-level concept — it’s an operational one. Every HTTP call you add to “share” a piece of state is a new failure mode at runtime. The button could have rendered based on a local date calculation that couldn’t fail. Instead, its existence now depended on a network call to a service designed for asynchronous message processing. If the consumer was down or slow, the button wouldn’t render.</p>
<p>When you DRY two pieces of code together, you’re not just removing duplication — you’re <strong>adding an edge to the dependency graph of your system.</strong> That edge has its own cost in flexibility, coordination, decomposition difficulty, and runtime availability. The cost is often higher than the cost of the duplication. But it’s invisible until the day you need to move one of those nodes, or one of them goes down at 3 AM.</p>
<p>Here’s the conversation worth having. Next time someone wants to extract:</p>
<blockquote>
<p>“What happens the first time these two callers need to behave differently? Right now they look the same, but they live in two different domains with two different futures. The moment the booking side needs a new currency rule and the reporting side doesn’t, we’ll add a flag. Then a second flag. Then a third caller shows up. Two years from now this is a method with seven parameters that nobody can read, and changing any one of them risks breaking five things. Or we leave it duplicated, and when the booking side needs a new rule, it changes one file. Which version of the codebase do you want to maintain?”</p>
</blockquote>
<p>The argument the script makes is one that doesn’t go away when you change where the code physically lives: <strong>shared code forces shared shape.</strong> Two callers that use the same function are constrained to keep behaving the same way, <em>or</em> the function grows parameters to accommodate the differences. That cost lives in the design, not in the build pipeline. A monorepo doesn’t fix it. It just makes it easier not to notice.</p>
<p>The math isn’t always this clean. The act of doing it changes the conversation from “good practice vs lazy” to “this cost vs that cost.” Which is the conversation you actually want to be having.</p>
<h2>Stable Contracts: When Slow Is the Point</h2>
<p>This is the post’s most counterintuitive argument, and probably the most important one.</p>
<p>When engineers argue against splitting a monolith, the argument often takes this form: <em>“If we split, we’ll need a shared library, and that’ll slow us down.”</em> True. And: <strong>that slowness is often what you want.</strong></p>
<p>A contract that moves slowly between two systems is a <em>feature</em>. The contract is what stops a change on one side from breaking the other by accident. The slowness is <em>intentional friction</em> — and intentional friction is the mechanism by which the contract gets respected. The trouble is that “removing friction” sounds, in every engineering culture I’ve ever worked in, like an unambiguous win. It isn’t. Some friction is the architecture doing its job.</p>
<p>Let me tell you about the BFF that stopped being a BFF.</p>
<p>I worked on a system with a backend and three clients — web, Android, and iOS — sitting behind a BFF (backend-for-frontend). Standard pattern: the BFF’s job is to maintain a stable contract for the clients while the backend evolves at its own pace. The mobile apps especially needed it; iOS and Android both have long-tail version distributions, and any breaking change in the backend’s payload propagates to every app version still in the wild.</p>
<p>Every backend change required a corresponding BFF change. That’s how the system was designed. It also felt — to the engineers shipping every day — like an obvious inefficiency. <em>“We’re touching two repos for one change.”</em></p>
<p>Someone came up with what sounded like a great idea: publish the backend’s schema at runtime and have the BFF dynamically map its response types from the live backend schema. <em>“Now we only have to update the backend. The BFF picks up the change automatically.”</em> Fewer pull requests. Fewer commits. Velocity up.</p>
<p>Ignoring the fact that the BFF’s literal job is <em>to maintain a stable contract</em>.</p>
<p>A few months after the project shipped, one of the backend engineers showed me a report he was proud of. It mapped which iOS versions in the wild were using which fields of his backend contract. He was excited — it would let him plan field deprecations more easily. <em>“I can see when nobody’s using the field anymore.”</em></p>
<p>My jaw dropped.</p>
<p>That report should never have existed. The backend engineer should not have been able to see iOS versions at all — that’s exactly what the BFF was supposed to hide. The friction the team had “removed” was the friction that decoupled the iOS release cycle from the backend release cycle. The BFF was still there in the codebase. Architecturally, it had ceased to exist. The backend engineer now had end-to-end visibility into the mobile clients, and was happily optimising past a boundary that no longer functioned as one.</p>
<p>We hadn’t sped anything up. We’d merged a four-system architecture into one tightly-coupled blob with extra hops. The deprecation report wasn’t a triumph; it was the failure mode showing itself in the most well-intentioned way possible. The team had built a tool to make the wrong work easier.</p>
<p>The principle:</p>
<blockquote>
<p><em><strong>The right amount of friction between two systems is whatever makes the cost of coupling them visible at the moment they’re being coupled.</strong></em></p>
</blockquote>
<p>Too little friction → developers couple things accidentally and discover the cost later (the BFF story, and most monorepos in my experience). Too much friction → nobody integrates anything and you have a coordination crisis. The right friction is <em>intentional</em> — chosen, named, defended — and it’s the friction that prompts a conversation.</p>
<p>Linus Torvalds has been shouting <em>“WE DO NOT BREAK USERSPACE!”</em> at kernel maintainers for thirty years. That’s a stable contract. It costs the kernel team enormous ongoing energy. They pay it on purpose. The slowness is what makes the kernel a thing the rest of the world can build on. The contract isn’t a bug. It’s the whole point.</p>
<p>If you’ve ever read <a href="https://blog.dicko.dev/posts/6-golden-rules-for-library-development-ensuring-stability-and-reliability/">6 Golden Rules for Library Development</a>, this is where it cashes out — the discipline of non-breaking changes is what <em>makes</em> slow contracts liveable. Without it, the slowness is just annoying; with it, the slowness is protective.</p>
<h2>When Sharing Is Actually Right</h2>
<p>The post can’t just be “stop sharing things.” That would be as bad as “share everything.” So here’s how to read the cases where sharing is the right call.</p>
<p><strong>Strong candidates for sharing</strong> — places where the signal is genuine, unifying knowledge:</p>
<ul>
<li><p><strong>Domain primitives that are foundational to the business.</strong> A Money type. A BookingId. A Currency enum. These are <em>ubiquitous language</em>, in Eric Evans' sense. They don't change often, and when they do, every consumer <em>should</em> see the change.- <strong>Cross-cutting infrastructure with a stable, narrow contract.</strong> Logging, tracing, auth. But — and this matters — the shared thing is the <em>interface</em>, not the implementation. See <a href="https://blog.dicko.dev/posts/managing-dependency-conflicts-in-library-development-advanced-versioning-strateg/">Managing Dependency Conflicts in Library Development</a> for the abstractions-library pattern.- <strong>Encoded business rules that are legally or contractually identical everywhere.</strong> Tax calculation. Regulatory compliance. PII redaction. If the business says <em>“this is the rule,”</em> there’s exactly one rule.
<strong>Weak candidates for sharing</strong> — places where the signal is accidental similarity:</p>
</li>
<li><p><strong>Anything that “looks similar today” but is owned by different teams.</strong> Different owners → different velocities → different futures. The similarity is a snapshot, not a structure.- <strong>Anything that crosses a service or domain boundary you’re actively trying to enforce.</strong> Sharing across the boundary undoes the boundary.- <strong>Anything that’s been DRY-ed up and immediately needed a flag parameter.</strong> That’s the signal that you’ve coupled two pieces of knowledge that wanted to diverge.
The question that does the work isn’t binary. It’s calibrated:</p>
</li>
</ul>
<blockquote>
<p><em><strong>“Given this code, this team, this service boundary, and this rate of change — where on the DRY dimension does this belong, and how confident am I that it belongs there?”</strong></em></p>
</blockquote>
<p>A Money type owned by one platform team across a stable business: high on the dial. The expected lifetime is years, the consumers are many, the contract is stable, the knowledge is genuinely singular. Extract it.</p>
<p>A “shared helper” between two teams that just split out of a monolith and are still discovering their own shape: low on the dial. The expected lifetime of the similarity is months, the consumers are diverging, and the cost of locking them together now is the cost of preventing the divergence that was the whole point of splitting. Leave it duplicated.</p>
<p>The same code in the same team six years apart can call for different answers. That’s not a bug in the framework. That’s the framework.</p>
<h2>A Sidebar: The Look of Shock</h2>
<p>I’ve told teams, very casually, <em>“just copy-paste it into the other repo.”</em> You’d think I’d stood on their cat. <em>“Copy. Paste?”</em> The look of shock — like I was some kind of heretic.</p>
<p>That reaction is the whole problem. DRY has stopped being a heuristic for these engineers and become a <em>moral framework</em>. Copy-pasting isn’t a technical choice with trade-offs — it’s <em>wrong</em>, in the same register as not writing tests. Which means the conversation can’t happen on technical grounds, because the engineer isn’t on technical grounds. They’re defending a value.</p>
<p>The job of the senior engineer in that moment isn’t to win the technical argument. It’s to give the team <em>permission to think technically</em> — to name the value, separate it from the heuristic, and ask the actual question: <em>what’s the cost of duplicating these eight lines versus the cost of coupling these two services?</em> Once the question is on the table, most engineers can answer it.</p>
<p>It’s the <em>asking</em> that’s hard.</p>
<h2>Too Much Doing, Not Enough Thinking</h2>
<p>The real villain in this post isn’t DRY. It isn’t the engineer who reflexively extracts a helper. It isn’t the reviewer who demands deduplication on every PR. The villain is <strong>the absence of the conversation that should have happened before the keystroke</strong>.</p>
<p>The engineer who reflexively DRYs a two-line duplication isn’t being dogmatic — they’re being <em>under-equipped</em>. They never had the five-minute conversation about coupling cost. Nobody walked them through the rule of three. Nobody showed them Metz’s essay. They learned DRY as a rule from a textbook, and nobody taught them the second half.</p>
<p>The God Helpers, the Christmas-tree parameter lists, the BaseControllers with 30 constructor parameters, the platform libraries that became their own bottleneck — none of these are caused by bad engineers. They’re caused by <em>too much doing, and not enough thinking</em>.</p>
<blockquote>
<p>“Premature optimization is the root of all evil.”* — Donald Knuth (paraphrasing Tony Hoare)*</p>
</blockquote>
<p>The analogue holds: <strong>premature abstraction is the root of an awful lot of system-level evil.</strong> Both are caused by the same instinct — solving for an imagined future rather than the actual present. Both look like good engineering at the moment they happen. Both pay back over years.</p>
<p>The fix is unglamorous. It’s a conversation between two engineers with experience, <em>before</em> the extract-method shortcut gets used. Five minutes of <em>“should we?”</em> prevents five years of <em>“how do we untangle this?”</em> And the way you make those conversations cheap and frequent is by sharing knowledge — workshops, brown bags, design reviews that actually review design, senior engineers who teach by walking through trade-offs instead of just shipping the answer.</p>
<p>I run workshops constantly — internally at work and in the wider engineering community. The reason isn’t that workshops are a goal. It’s that they’re the cheapest mechanism humans have ever invented for the conversation that prevents the wrong abstraction. Five engineers in a room, an hour of their time, one shared mental model of why coupling has a cost. The ROI is enormous, and it’s underrated because the bug it prevents is the one that never gets written.</p>
<p>This is also the inverse of something I wrote about in <a href="https://blog.dicko.dev/posts/code-entropy-the-silent-killer-of-engineering-velocity/">Code Entropy: The Silent Killer of Engineering Velocity</a> — copy-paste <em>coding</em> is an entropy source, yes. But copy-paste <em>avoidance</em>, taken to its extreme, is also an entropy source. Just a slower-acting one. Both are symptoms of the same disease: applying a rule without the conversation that turns it into judgement.</p>
<p>DRY failed because we made it a rule.</p>
<p>It isn’t a rule. It’s a dimension, and the engineering skill is knowing how far in either direction your current context allows you to go.</p>
<p>There is no such thing as best practice. There is only adequate practice given the current context.</p>
<p>Now, if you’ll excuse me, I need to go review a pull request. Someone on my team has extracted a helper that takes seven parameters, and I’m about to type a polite comment asking whether this helper is really one function or three pretending to be one. I’ll probably get it wrong the first time. I usually do. But the conversation is the point.</p>
]]></content:encoded>
  </item>
  <item>
    <title>The Floor Plan Is Your First Architecture Decision</title>
    <link>https://blog.dicko.dev/posts/the-floor-plan-is-your-first-architecture-decision/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/the-floor-plan-is-your-first-architecture-decision/</guid>
    <pubDate>Mon, 04 May 2026 08:01:40 GMT</pubDate>
    <category>engineering-leadership</category>
    <category>remote-work</category>
    <category>engineering-culture</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <description>Or: Your Seating Chart Is Picking Your Software Architecture, Whether You Like It or Not</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/the-floor-plan-is-your-first-architecture-decision/cover.webp" alt=""></figure>
<p>It was 10:47 on a Tuesday when I noticed her. A junior engineer, two desks down from me, staring at Slack with the particular stillness of someone willing a green dot to turn into a reply. Her Thai iced tea was sweating onto the desk next to her elbow. I asked what she was working on. She told me she was waiting on the team in the row behind us — a team I could see over her shoulder, headphones on, less than ten metres away.</p>
<p>That gap — between <em>the people she needed</em> and <em>the channel she was using to reach them</em> — is the entire post. It’s also one of the most expensive design decisions your organisation has made, and almost nobody treats it as a design decision at all.</p>
<h2>The Decision Nobody Reviews</h2>
<p>Most architecture decisions get a design review. They get RFCs. They get whiteboard sessions where someone draws boxes and someone else asks pointed questions about consistency guarantees. Then there’s the seating chart, which gets handed to someone in facilities with a spreadsheet of headcount and a cheerful request to “fit everyone in.”</p>
<p>That decision — who sits next to whom, who shares a kitchen, who’s within earshot of whom — will shape the information flow of your engineering organisation more than any service mesh, any internal wiki drive, any Slack channel naming convention. It is, in the strict sense, an architecture decision. It determines what is coupled to what, what information flows where, and what conversations happen at all. And we’re outsourcing it to a spreadsheet.</p>
<p>This post isn’t going to relitigate remote vs in-office. There’s a <a href="https://blog.dicko.dev/posts/co-location-remote-first-and-the-rollback-every-few-decades/">companion post</a> for that, and there’s plenty of noise on it elsewhere. The argument here is narrower and a lot less contested: <strong>wherever and whenever your engineers are physically together, the floor plan they sit in is itself a design decision.</strong> Pick the couplings you want. Pick the ones you’re willing to break. Treat the seating chart like architecture, because it is.</p>
<blockquote>
<p>“We shape our buildings; thereafter they shape us.”* — Winston Churchill, House of Commons rebuilding debate, 28 October 1943*</p>
</blockquote>
<p>Churchill was arguing for rebuilding the bombed Commons chamber in its original cramped, adversarial shape — because the shape, he said, had produced the British two-party system. The chamber was too small to hold all the MPs. The benches faced each other. You couldn’t slip in unnoticed; you had to physically cross the floor to change sides. The architecture was producing the politics. The same logic applies to your engineering floor. You picked a layout. The layout is now picking how your engineers work.</p>
<h2>The Snacks Engineer</h2>
<p>A few months ago I was talking to one of our engineers about a team he kept running into for code reviews. The team was always slammed — heads-down work, calendars full, the usual. But this engineer’s PRs kept getting merged faster than anyone else’s. So I asked him how.</p>
<p>“I just walk over,” he said. “Sometimes they’re not free, so I wait at a desk nearby until they are.”</p>
<p>I was a little surprised. “Don’t they find that intimidating? You just sitting there?”</p>
<p>He grinned. “No — because I bring them snacks. I’m nice about it.”</p>
<p>That’s the entire post in one anecdote. Not a productivity hack. Not a process. A guy with a packet of orange Kit Kats from someone’s last Japan trip, walking over, waiting politely, getting his PR reviewed in twenty minutes when the Slack version would have been a day. Multiply that by a thousand interactions a year across a hundred engineers, and you’ve quietly redesigned how your organisation moves.</p>
<p>The Slack version isn’t slower because Slack is bad. Slack is fine. Teams is, well, less fine — but you knew that. Whatever your tool of choice is, that’s not the problem. The problem is that <strong>asynchronous communication has a built-in tax that synchronous communication doesn’t</strong>, and that tax compounds across every interaction in every day in every team. Walk over once and you save twenty minutes. Walk over a hundred times — politely, with snacks — and you’ve outpaced the team that’s still typing.</p>
<h2>The Allen Curve, or: Why “Just Slack Them” Is a Lie</h2>
<p>In the late 1970s, an MIT professor named Thomas J. Allen did the experiment most engineering leaders have never heard of. He measured how often engineers communicated as a function of how far apart their desks were. The result was an exponential decay: engineers two metres apart communicated four times more often than engineers twenty metres apart. At fifty metres, weekly communication had effectively collapsed. On a different floor or in a different building, they almost never communicated at all.</p>
<p>The natural objection is the obvious one — and it’s the assumption this post is pushing back on: <em>“That was the 1970s. We have Slack now. Email. Teams. Zoom. The internet. The tools have replaced presence.”</em></p>
<p>They have not. Allen ran follow-ups in the 2000s and the curve still held — and crucially, <em>all</em> communication channels decayed with distance, not just face-to-face. Phone, email, video — all of them. The richer your in-person relationship, the more you also use the digital channels. Out of sight, out of mind, on every channel, not just the in-person one. The tools didn’t replace proximity; they amplified the relationships that proximity already produced. Your most-DM’d colleagues are also the ones you sat near, ate with, ran into at the kitchen. Take the proximity away, and the digital channel slowly atrophies too.</p>
<p>A 2013 University of Michigan study put numbers on the modern version. Researchers in the same building were 33% more likely to collaborate than those in different buildings. Researchers on the same floor were 57% more likely than those on different floors. Same internet. Same email. Same video tools. The floor plan is still doing the work.</p>
<p>This is the part that’s hard for engineering leaders to internalise: <strong>the seating chart builds a second org chart, running in parallel to the official one.</strong> Your reporting lines are still the reporting lines — they decide who you escalate to, who reviews your performance, who owns the team’s outcomes. But your real day-to-day communication paths are a function of where people sit, not of who technically reports to whom. Conway’s Law says your software will mirror your communication structure. Your communication structure is the <em>parallel</em> org chart, the one you didn’t draw, the one that runs on proximity. So the floor plan is, transitively, helping to design your software.</p>
<p>Which is why <strong>treating the seating chart as a logistics question is the quiet failure most engineering organisations make.</strong> When facilities sends the seat map for the new floor and a manager forwards it to the team for “any concerns,” that’s not a process — that’s outsourcing the architecture. The tools-replace-presence assumption is what makes that outsourcing feel safe: if Slack is doing the real work, who cares where the desks go? But the data has been clear since the late 1970s. The desks are doing the real work. The tools are a top-up.</p>
<h2>What Co-location Actually Buys You</h2>
<p>Be concrete for a second about what you lose when you lose proximity. The traditional phrase is “water cooler conversations” — fine, but that understates it. Here’s a less polite version of the list.</p>
<figure><img src="https://blog.dicko.dev/posts/the-floor-plan-is-your-first-architecture-decision/image_2.webp" alt="" loading="lazy"></figure>
<p>The pattern: <strong>proximity creates a steady drip of low-cost information flow that nobody planned, nobody scheduled, and nobody can replicate in a tool.</strong> The Slack equivalent of “overhearing the next team’s standup” is reading their channel — except nobody actually reads other teams’ channels, because there are forty of them and the signal-to-noise is dreadful. So the information doesn’t transfer. The team next door ships something you should have known about, and you find out three weeks later in a stakeholder meeting where someone says it like it’s old news.</p>
<p>The Gemba walk deserves its own paragraph. The Lean tradition uses the term for factory floors: go to the place where the work happens, observe, ask. In an engineering org, it’s the senior engineer or manager who walks up to people’s desks and just asks: <em>what are you working on, and what’s blocking you?</em> The information you get from fifteen minutes of that is information you would never get from a status report.</p>
<p>I do this regularly, and the pattern is consistent. I’ve sat down with engineers and uncovered issues with testing frameworks that were hard to use and slowing the team down — issues that had never been escalated. I’ve found patterns of last-minute requests from supporting teams that nobody had thought to raise as a systemic problem. The list goes on. None of these would have come up in a standup. Standups are about what you’re doing today, not about the recurring friction that’s been quietly costing you a day a week for six months. None of these would have come up in a retro either, because retros tend to surface the <em>acute</em> — the thing that broke this sprint — rather than the <em>chronic</em> — the thing that’s always slightly broken and that everyone has learned to work around.</p>
<p>You can’t really do this on Teams. Or you can, but it lands differently. A drive-by question at someone’s desk is conversational. The same question in a DM is an interrogation.</p>
<p>The corollary to the Gemba walk is the kitchen test. When I haven’t seen someone in a while, I always lead with: <em>working on anything interesting lately?</em> It sounds like small talk. It’s not. The question gives them permission to talk about what they actually care about — which is almost always also what they’re stuck on, what they’re proud of, or what they wish someone else knew about. Many times in large organisations, two teams are solving the same problem in isolation, and neither knows the other exists. The kitchen conversation surfaces those overlaps more often than not.</p>
<p>There is no Slack channel for “interesting things people are working on that you should know about.” There can’t be. The moment you create one, it becomes another channel nobody reads. The kitchen, for as long as engineers go there, is the channel.</p>
<h2>The Pixar Principle: Designing for Collisions</h2>
<p>Some leaders intuited all of this before the data caught up. Steve Jobs is the canonical example, and the Pixar building is the canonical artefact.</p>
<p>When Pixar moved to its Emeryville campus in the late 1990s, the original design separated computer scientists, animators, and editors into three buildings. Jobs killed it. He insisted on one building, with a giant central atrium — and then, infamously, put the cafeteria, the mailboxes, the meeting rooms, and (most insidiously) the only bathrooms inside that atrium. People had to come to the centre. People had to bump into each other. John Lasseter’s verdict afterwards was unequivocal: he’d never seen a building that produced collaboration like that one.</p>
<p>The design principle: <strong>proximity by accident is unreliable; proximity by design is leverage.</strong> Jobs didn’t trust serendipity. He engineered the floor plan so that serendipity was the default outcome.</p>
<p>For most of us, the takeaway isn’t “build an atrium.” We aren’t designing buildings. The takeaway is the disposition: <strong>the seating chart is a forcing function. Use it.</strong> Two teams that need to integrate? Sit them together. A platform team that’s drifting from its consumers? Put them next to a consumer team. A new hire who needs to absorb culture? Don’t put them on the empty desk in the corner — put them in the middle of the noisiest, most opinionated team you have.</p>
<p>There’s a related insight from a different field. Jane Jacobs, writing about cities in 1961, argued that street safety came from “eyes on the street” — the casual, ambient awareness of residents and shopkeepers going about their daily business. The principle generalises beyond cities: <strong>public visibility is its own coordinating mechanism.</strong> When you can see what’s happening, you act on what you see. When you can’t, you don’t.</p>
<p>In an engineering office, eyes on the street is the manager who notices a junior engineer staring at the same stack trace for forty minutes and walks over. It’s the senior who sees a whiteboard sketch on the way back from lunch and goes <em>wait, that’s wrong, here’s why</em>. It’s the architect-equivalent — at Agoda we don’t have architects, just engineers, but the same role exists informally — who walks past a cluster of people arguing and joins the conversation uninvited because the conversation is important. None of this happens on Slack, because Slack hides everything that isn’t explicitly shared. The kitchen, the desk pod, the open floor — those are streets. Eyes on them produce coordination at no cost.</p>
<p>Our level 6 in Bangkok is not Pixar’s atrium, and I’m not going to pretend it is. But the mechanic is the same: when teams sit close enough that walking over is cheaper than messaging, the conversations that matter happen on the way to the kitchen instead of three days later in a calendar invite.</p>
<h2>The Loud Team and the Quiet Team</h2>
<p>A few years ago I had a team that was very collaborative. They worked out loud — and I mean genuinely out loud, not metaphorically. Throughout the day you’d hear them. Not the whole team at once, but at any given moment at least two of them would be pairing for an hour, or talking through the current piece of work to align on approach, or sorting out a merge conflict at someone’s desk. They were all working towards the same goal, and the conversations were about the work. The talking <em>was</em> the working.</p>
<p>When we shuffled the floor plan, they ended up next to a team from a different department who happened to be working on a closely related problem.</p>
<p>The other team had a completely different working style. They sat quietly at their desks all day, headphones on, world shut out. Both teams were good. Both teams shipped. But the seating arrangement produced a specific kind of complaint — the quiet team kept telling me the loud team was disturbing them. Couldn’t focus. Too noisy. Could we move them?</p>
<p>We didn’t move them. The loud team won.</p>
<p>That sounds harsh, so let me explain why. <strong>Software engineering is a collaborative endeavour.</strong> If you’re sitting on your own all day with headphones on, deeply focused, you are also building in isolation. Less feedback. Fewer eyes. Less informal pressure-testing of your assumptions. The code that comes out the other end is more likely to have blind spots — not because the engineers are worse, but because the working style produces a worse coverage pattern. The collaborative team was, by accident of how they worked, doing constant peer review just by being audible to each other. This connects to a thread we’ve pulled on before in <a href="https://medium.com/beer-and-servers-dont-mix">Code Entropy: The Silent Killer of Engineering Velocity</a> — entropy grows fastest where there are fewer eyes on it.</p>
<p>What happened next is the part I didn’t expect. Over time — not weeks, more like a quarter — the quiet team caught on. They started picking up some of the louder team’s habits. Talking through problems. Pulling each other into conversations. Working with at least one ear on the room rather than two ears in a hoodie. They didn’t become a different team. They just became a team with more open channels. And their work got better. The complaints stopped not because the loud team got quieter, but because the quiet team realised they’d been missing something.</p>
<p>The lesson isn’t that loud teams are better or quiet teams are wrong. Some work genuinely needs deep focus, and noise-cancelling headphones are a real productivity tool. The lesson is that <strong>the seating arrangement made a working style visible that hadn’t been visible before, and the comparison taught the quieter team something they couldn’t have learned in isolation.</strong> Sometimes the floor plan should be deliberately uncomfortable, because the discomfort is where the upgrade happens.</p>
<h2>When You Can’t Co-locate Every Day, Co-locate Hard for a Week</h2>
<p>The point of this post isn’t that everyone should be in the office every day. That argument lives in <a href="https://blog.dicko.dev/posts/co-location-remote-first-and-the-rollback-every-few-decades/">the companion post</a> and isn’t worth relitigating here. The more interesting question is the constructive one: when full-time co-location isn’t on the table — distributed teams, external partners, hybrid policies, four offices in three countries — what do you do?</p>
<p>The answer we’ve found is the <strong>acceleration week</strong>: a short, deliberate, intensive burst of full co-location aimed at a specific problem. Not an offsite. Not a hackathon. A working week with people who normally don’t share a floor, in the same room, on the same problem, with all the boring daily structure that real work requires.</p>
<p>A recent one. We brought eight engineers from three external companies to Bangkok to work alongside fifteen of our own — twenty-five people total, five days, in the same room. The trigger was a serious production incident that had landed earlier in the week, in exactly the area the visiting engineers were already coming to work on. By Friday, the legacy architecture behind the incident had been replaced with something safer, and a stack of long-pending merge requests had finally landed.</p>
<p>The numbers aren’t the point. The <em>rate</em> is. One of the engineering managers put it cleanly afterwards: <em>if we hadn’t done this, the migration probably would have lasted a year. All of those problems we solved right away. Otherwise that becomes communication on Slack, and a reply maybe next week, because we’re busy or they’re busy.</em></p>
<p>That’s the entire economic argument for acceleration weeks in one sentence. A year of intermittent Slack-and-meeting work, compressed into five days of focused, high-bandwidth collaboration — because the cost of asking a question dropped to zero and the cost of context-switching dropped with it. A senior engineer described what made it different: <em>normally, day to day, I’m doing ten things at a time and nothing gets done. This kind of stuff is really good productivity.</em></p>
<p>A few mechanics worth naming, because the format is reproducible:</p>
<p>Three stand-ups a day, not one. When work moves at compressed-time speed, you need to re-sync more often. Twice a day at minimum, three when things are flying. Catered meals, deliberately — people stay together, lunch is not a break from the work, it’s where half the cross-pollination happens. The kitchen test, applied to a hotel ballroom. Pre-share the context two weeks ahead — architecture overviews, tooling guides, reading material. The first day shouldn’t be onboarding. Invite the supporting teams from day one — infra, security, on-call — because the bottleneck on day three is always the team you didn’t invite. And accept productive discomfort. One of the visiting engineers nailed it: <em>you kind of have to be comfortable being uncomfortable. The looseness around the planning is why you’re able to find the problem, get the people, and tackle it — because there’s no set scope.</em> Tightly-scoped weeks produce tightly-scoped outputs. The good ones leave room for the problem to define itself.</p>
<p>The most telling quote came from one of our own engineers. Not pitching the format. Just describing what it felt like:</p>
<blockquote>
<p>This is more what it used to be like in Agoda when I first got here, because everyone was in the office. Everyone in our department was on level six, and we only took up half the level, so you could see everyone from one end. It’s much easier to get stuff done. If you’re waiting for a code review, you wouldn’t use the PR bot to ping someone on Slack. You’d go and sit down at the desk and say ‘hey, can you do this code review for me? I’m waiting.’</p>
</blockquote>
<p>The acceleration week didn’t invent a new mode of working. It rebuilt, briefly, the working environment that produced casual collaboration in the first place — and reminded everyone in the room what that mode actually feels like. <strong>Proximity is a tool, and tools can be used at different cadences.</strong> When the floor plan can give you proximity full-time, design it well. When it can’t, design the intensive that gives you proximity in a burst. The mechanic underneath is the same — put the people in a room, make asking a question free, watch what happens to the rate of progress.</p>
<h2>The Generational Gap</h2>
<p>There’s a careful version of the next point and a lazy version. The lazy version is <em>kids these days don’t know how to work in an office</em>, which is wrong, condescending, and unhelpful. The accurate version is harder to say but worth saying: <strong>this is a real generational gap, and it goes deeper than office culture.</strong></p>
<p>Engineers who started their careers in 2020, 2021, or 2022 worked entirely remote during the formative period when most people pick up the unwritten rules of how work works in person. They didn’t get to watch a senior engineer push back from the desk, walk twenty metres, and end a Slack thread that had been spinning for three hours. They didn’t get to see how someone reads the room before interrupting. They didn’t develop the muscle for the bounded, polite, in-person ask. None of this is a character failing. It’s a missing chapter in their working education, and it has to be backfilled the way any missing chapter does — by being taught.</p>
<p>Which brings me back to the engineer at the start of this post.</p>
<p>I stopped by and asked what she was working on. She told me she was waiting on a reply. I looked over at the team — same room, same floor, less than ten metres. I asked her: <em>why don’t you just go talk to them?</em></p>
<p>Her answer was kind, not lazy. <em>I don’t want to bother them.</em></p>
<p>That sentence captures the gap in one line. She didn’t think she’d be turned away. She didn’t think they’d be rude. She just hadn’t internalised that walking over isn’t an imposition — it’s the <em>normal</em> way engineers work in person, and it’s almost always faster, lower-friction, and counterintuitively <em>less</em> disruptive than a Slack message that pings someone out of focus mid-task. A walk-over is bounded: it ends when the conversation ends. A chat thread is open-ended — it pings, gets ignored, gets replied to two hours later, gets clarified, gets clarified again, and consumes everyone’s attention in fragments.</p>
<p>I told her — gently, because the fear of bothering people is genuine and considerate, not a fault — that for this kind of question, walking over is the right move. We talked through how to do it. Catch their eye first. Ask if it’s a good moment. Keep it short. Leave when it’s done. The skills of the walk-over. None of them are obvious. All of them have to be modelled.</p>
<p>For leaders, this is a teaching problem. Engineers who started remote will not magically learn office culture by being put in an office, any more than a kid who missed two years of in-person school catches up just by being back in a classroom. They have to be coached, deliberately, by the senior engineers who do remember the rhythms. That coaching is a real responsibility, and it’s not optional. The skill is real, it can be taught, and it has to be taught — because the alternative is a workforce that never developed the muscle and an organisation that never figures out why its best ideas don’t propagate.</p>
<blockquote>
<p>“The single biggest problem in communication is the illusion that it has taken place.”* — George Bernard Shaw*</p>
</blockquote>
<p>Shaw was being witty about Edwardian dinner parties. He could have been describing the Slack thread that ended in <em>thanks, makes sense</em> when neither party actually understood the other.</p>
<h2>The Manager’s Floor Plan Playbook</h2>
<p>A few opinions, sharpened.</p>
<p><strong>Sit teams together that need to integrate.</strong> The frontend team and the backend team for the same product. The platform team and its biggest consumer. The two teams whose services have the most cross-team API calls. If they need to talk every day, stop making them schedule it.</p>
<p><strong>Sit teams <em>apart</em> that you want to evolve independently.</strong> This is the less-obvious application. If two teams need to make different architectural choices, putting them together can produce harmful coupling. They start adopting each other’s conventions, and the system loses the diversity that lets it experiment. Sometimes distance is what you want.</p>
<p><strong>Walk the floor.</strong> Gemba is not a workshop. It’s a habit. Twice a day, walk through the area where your engineers sit. Don’t aim for anyone. Let yourself notice what’s happening. Ask what people are working on. The information you’ll get in fifteen minutes of walking is information you would never get from a status report.</p>
<p><strong>Treat seating changes as architecture reviews.</strong> When facilities sends you the seating chart for the new floor, don’t approve it as a logistics question. Read it as a system design. What does this proximity create? What does it break? What’s the new Conway’s-Law-implied software architecture? If you can’t answer those questions, you’re outsourcing your architecture to facilities.</p>
<p><strong>Build for collisions, not for quiet.</strong> Open-plan offices are unpopular for good reasons — they’re noisy, they break focus work, they cost more in headphones than they save in real estate. But the collision benefit is real. The right design is hybrid: focused desk pods for deep work, plus shared spaces — kitchens, big monitors on walls, sofas, whiteboards — where collisions happen. Pixar’s atrium is the maximalist version. The minimalist version is <em>make sure two teams have to walk past each other to get coffee.</em></p>
<p><strong>Coach the walk-over.</strong> For engineers who started remote, the skill of going to someone’s desk doesn’t exist yet. Name it. Model it. Tell them, when they’re stuck waiting on a Slack reply: <em>go talk to them. Bring snacks if you’re worried about being a bother.</em> Make it explicit that this isn’t bothering people. This is how the work is actually supposed to flow.</p>
<h2>The Question That Cuts Through Everything</h2>
<p>When you’re deciding where someone should sit, ask yourself one question:</p>
<blockquote>
<p><em><strong>What conversations do I want to happen by accident?</strong></em></p>
</blockquote>
<p>Then put those people near each other. The conversations you have to schedule are the conversations that are too expensive to happen often. The conversations that happen by accident are the ones that compound. Pick which is which, and seat accordingly.</p>
<blockquote>
<p>“There must be eyes upon the street, eyes belonging to those we might call the natural proprietors of the street.”* — Jane Jacobs, <em>The Death and Life of Great American Cities</em> (1961)*</p>
</blockquote>
<p>Jacobs was writing about urban safety. The principle generalises to any environment where coordination matters and full surveillance isn’t an option. Visibility is the cheapest coordinating mechanism humans have ever invented. Engineers in line of sight of their colleagues coordinate by accident. Engineers behind walls and time zones coordinate by meeting. The first is free. The second is not.</p>
<p>Now, if you’ll excuse me, I need to go find an engineer on level 6 who’s been waiting two days for a code review from someone sitting fifteen metres away. He has Slacked them three times. They’re heads-down with headphones on. I’m going to walk him over there. I’ll bring snacks. I’m nice about it.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Design Thinking Is Not a Workshop — It’s a Hiring Criteria</title>
    <link>https://blog.dicko.dev/posts/design-thinking-is-not-a-workshop-its-a-hiring-criteria/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/design-thinking-is-not-a-workshop-its-a-hiring-criteria/</guid>
    <pubDate>Thu, 09 Apr 2026 06:59:12 GMT</pubDate>
    <category>product-management</category>
    <category>design-thinking</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>engineering-leadership</category>
    <description>Or: Why Your Two-Day Empathy Exercise Won’t Fix a Feature Factory</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/design-thinking-is-not-a-workshop-its-a-hiring-criteria/cover.webp" alt=""></figure>
<p>It was 10:15 on a Wednesday — mid-sprint planning — and Somchai was leaning forward in his chair, palms flat on the table, making his case for the third time. Not about scope. Not about user impact. About whether the ticket was a five or an eight. The meeting room on level 6 had that particular warmth of too many people and not enough ventilation, the glass walls fogging slightly at the edges, and the engineers around the table had developed a sudden fascination with their laptops. Someone’s Thai iced tea sat untouched, condensation rings spreading across the table. “But my points,” Somchai said — not loud, but with the kind of quiet intensity that makes everyone else go very still. Nobody in that room was thinking about the user. Least of all Somchai.</p>
<p>Here’s the thing: Somchai wasn’t broken. Somchai was perfectly calibrated to the system we’d built around him. Months of velocity-driven conversations had taught him that points were the scorecard, and he was playing to win. The fact that nobody — including Somchai — could tell you whether anything he’d shipped in the last quarter had moved a single metric? That wasn’t Somchai’s failure. That was ours.</p>
<p>There’s a two-day design thinking workshop happening right now at a company somewhere in the world. Forty engineers are learning about empathy maps. They’ll do a post-it note exercise. They’ll prototype something in cardboard. They’ll go back to their desks on Monday and build exactly what the product spec tells them to build. Nothing will change. And the reason nothing will change has nothing to do with the workshop’s content and everything to do with a fundamental misunderstanding: most companies treat design thinking as a training programme when it should be a hiring criteria.</p>
<h2>The Feature Factory Is a Hiring Problem</h2>
<p>John Cutler’s “<a href="https://cutlefish.substack.com/">12 Signs You’re Working in a Feature Factory</a>” resonated so widely because it described a symptom almost everyone recognises — teams cranking out features nobody measures, nobody validates, nobody questions. The standard diagnosis is that it’s a process problem. Fix it with OKRs. Add measurement rituals. Run product discovery sprints. And yes, those help. But they all share the same assumption: that the people executing the process have the underlying capability to think in outcomes rather than outputs.</p>
<p>What if they don’t?</p>
<p>Most engineering interviews test for technical competence — can you solve this algorithm, design this system, debug this code? They don’t test for product instinct — can you question whether this feature should exist, identify the user need behind the request, recognise when a solution is technically elegant but practically useless? The result: you hire engineers who are excellent at building things and terrible at questioning whether those things should be built. Then you send them to a workshop and wonder why nothing changes.</p>
<p>Here’s the number that should keep you awake: according to research by Ronny Kohavi at Microsoft, <a href="https://blog.dicko.dev/posts/i-dont-care-what-you-build-and-neither-should-you/">only 10–30% of product changes actually improve metrics</a>. That means 70–90% of what we build doesn’t move the needle. The engineer who asks “should we build this?” before writing a line of code is worth more than the engineer who builds it twice as fast. We just don’t interview for that engineer.</p>
<p>I’ve written before about <a href="https://blog.dicko.dev/posts/i-dont-care-what-you-build-and-neither-should-you/">how engineering leaders should focus on measurement over implementation</a> — my old CTO Yaron’s relentless “what’s the KPI?” approach. But Yaron’s question only works if the engineers in the room are capable of thinking in KPIs. If their mental model is “I ship story points,” no amount of KPI-focused leadership will bridge the gap. You’re layering methodology on top of a capability gap.</p>
<h2>What You’re Actually Looking For</h2>
<p>This isn’t a checklist. It’s a mental model for the traits that separate product-minded engineers from pure implementers.</p>
<p>The first trait is curiosity about the user, not just the technology. It’s the engineer who reads the Jira ticket and asks “why does the user need this?” before asking “how should I implement this?” This isn’t about being difficult — it’s about a natural orientation toward understanding the problem before jumping to the solution. When you describe a feature request to a candidate in an interview, do they ask about the user context, or do they jump straight to architecture?</p>
<p>The second is the instinct to question the spec. Not contrarianism — genuine critical thinking about whether the proposed solution matches the stated problem. The engineer who says “the user asked for a filter, but looking at the data, they’re really trying to find X — would a smarter default solve it better?” This connects directly to Amazon’s “working backwards” approach: starting from the customer need and working backward to the solution, rather than starting from a technical capability and looking for a problem to solve.</p>
<p>The third is comfort with ambiguity. Product problems are messy. The user doesn’t know what they want. The data is incomplete. The requirements are contradictory. Engineers who need a perfectly specified ticket before they can start are implementers, not product engineers. The ones who can operate in the grey space — who can take an imperfect understanding and still make progress — are the ones who build products that users actually want. I’ve written about how this same capability manifests as <a href="https://blog.dicko.dev/posts/your-engineers-cant-negotiate-and-its-your-fault/">constraint literacy</a> — the ability to negotiate the gap between what’s asked and what’s possible, rather than defaulting to hero mode or wall mode.</p>
<p>The fourth is business context awareness. Does the engineer understand <em>why</em> the company exists, <em>how</em> it makes money, and <em>where</em> this feature fits in that picture? An engineer who understands the business model will naturally question features that don’t connect to value creation. As I wrote in <a href="https://blog.dicko.dev/posts/semantic-monitoring-the-question-youre-not-asking-about-your-production-systems/">Semantic Monitoring</a>, the question “what does this application do?” separates engineers who own their systems from those who merely operate them. The product engineer answers with the problem their system solves. The implementer answers with the tech stack.</p>
<h2>How to Interview for It</h2>
<p>Most engineering interviews are 100% technical. Adding a single product-thinking question changes the signal dramatically.</p>
<p><strong>The “Why Would We Build This?” question.</strong> Present a real feature request — sanitised from your own backlog. Ask the candidate to evaluate it. Not to design the solution — to question whether it should be built. Good signals: they ask about user data, suggest alternatives, identify assumptions, propose a way to validate before building. Bad signals: they immediately start designing the solution, don’t ask about the user, treat the spec as given.</p>
<p><strong>The question I actually use: “Tell me at a high level what your application is and what it does?”</strong> It’s purposely broad and open-ended. Deceptively simple. The way candidates answer reveals how they understand their system. Some start with what the UI looks like. Some start with the technical architecture. None of these are wrong answers — but all of them tell you about the person.</p>
<p>The product engineers? They start with the problem their system is solving. “We help users find the right hotel for their trip by…” or “Our system processes payments across 300+ methods so that…” That’s the signal. That’s the person who understands <em>why</em> they’re building, not just <em>what</em> they’re building.</p>
<p>An engineer who describes their application as “a React frontend with a .NET backend that talks to three microservices” has told you their mental model is technical. An engineer who says “we help small businesses manage their inventory so they don’t lose sales to stockouts” has told you their mental model is user-centred. Both can write excellent code. Only one will question whether the feature they’re building actually solves a user problem.</p>
<p>This question works because it can’t be prepared for. There’s no “right” answer to study. It reveals a worldview, not a skill. And it’s something any interviewer can use — you don’t need to restructure your entire <a href="https://blog.dicko.dev/posts/the-hiring-bar-is-a-lie-you-tell-yourself/">hiring process</a> to add it.</p>
<h2>Why Workshops Fail (And What Works Instead)</h2>
<p>Workshops fail because they try to install a capability through information transfer. You can’t install empathy through a slide deck any more than you can install taste through a coding standard.</p>
<p>The basketball coach Pat Riley once said, “Excellence is the gradual result of always striving to do better.” He wasn’t talking about software, but he was describing exactly why a two-day workshop doesn’t work — excellence in user empathy, like excellence in anything else, comes from daily practice embedded in the work itself, not from a one-off event followed by business as usual.</p>
<p>So what actually works?</p>
<p><strong>Exposure to users — not user research reports.</strong> Engineers sitting in on user testing sessions. Not watching a summary — sitting there, watching a real person struggle with the thing they built. This is visceral. It can’t be abstracted away into a report. When an engineer watches a user fumble through a booking flow they built, the business impact becomes immediate and personal in a way no dashboard ever will. We use the Jobs to be Done framework to bring engineers into product thinking — understanding the “job” the user is hiring the product for. It’s not about making engineers into designers. It’s about developing the empathy muscle that makes product instinct possible.</p>
<p><strong>Pairing engineers with designers during discovery.</strong> Not the handoff model — <a href="https://blog.dicko.dev/posts/the-design-handoff-is-a-lie-why-your-engineers-and-designers-are-playing-telepho/">design produces spec, engineering implements</a>. The collaboration model — engineering and design explore the problem space together. When engineers participate in discovery, they develop the empathy that workshops try to teach. They surface constraints that change the solution. They understand the <em>why</em> behind the <em>what</em>, and that understanding compounds with every project.</p>
<p><strong>Celebrating “we didn’t build it” decisions.</strong> Culture eats process. If your organisation only celebrates shipping, engineers will optimise for shipping. If you also celebrate the decision <em>not</em> to build something — because the team determined the user didn’t need it, or found a simpler solution — you create an environment where product thinking is rewarded. Researchers at the University of Virginia, led by Leidy Klotz, found that people overwhelmingly <a href="https://blog.dicko.dev/posts/the-question-that-changes-everything/">default to adding rather than subtracting</a> when asked to improve something. The engineer who asks “should we build this?” is fighting a cognitive bias that’s hardwired into all of us. That deserves celebration, not a puzzled look from the Product Owner.</p>
<p><strong>Making “how will we know?” a reflex, not a ritual.</strong> When every engineer asks “how will we know if this works?” naturally — before writing a line of code — you’ve succeeded. When it requires a facilitator to prompt it, you haven’t. This is the daily practice that no workshop can replace.</p>
<h2>The Organisational Immune System</h2>
<p>Why do companies that know about design thinking, that send people to workshops, that talk about user-centricity, still operate as feature factories? Because the organisational immune system rejects the change.</p>
<p>“We don’t have time for discovery.” Same energy as “we don’t have time to do it right.” The cost of not doing discovery is invisible until a feature ships and nobody uses it — which, remember, happens 70–90% of the time.</p>
<p>“That’s the PM’s job.” The division of labour that creates the feature factory: PMs decide what, engineers decide how. In a product engineering model, engineers participate in the <em>what</em> — not to overrule PMs, but to bring technical constraints and possibilities into the conversation earlier. This is <a href="https://blog.dicko.dev/posts/teaching-engineers-to-speak-in-constraints/">constraint-talk in action</a> — and it requires product instinct to do well.</p>
<p>“Our interview process is standardised.” Adding a product-thinking question feels risky because it’s “subjective.” But so is every system design question — we just pretend it isn’t. The interviewer’s judgment is already the mechanism. You’re just adding a dimension to what they’re judging.</p>
<p>“We hire for technical excellence and train for everything else.” This is the assumption this entire post challenges. You can teach someone React. You can teach someone system design. You cannot teach someone to care about users. You can create environments where caring is expected and rewarded — but the seed has to be there. As Tim Brown of IDEO argued in <em>Change by Design</em>, what you’re looking for is the “T-shaped” quality — deep technical skill with a genuine, intrinsic curiosity about the people who use what they build.</p>
<p>There’s a cultural dimension here worth naming. In a South East Asian engineering context — the same dynamic I explored in <a href="https://blog.dicko.dev/posts/teaching-engineers-to-speak-in-constraints/">Teaching Engineers to Speak in Constraints</a> applies to product instinct: engineers may not question the spec not because they lack product thinking but because culturally they don’t feel it’s their place. In these contexts, screening for product instinct at hiring is even more important, because the environment may not naturally encourage it to emerge. You need to select for the seed and then build the culture that gives it permission to grow.</p>
<h2>The Transition: From Feature Factory to Product Engineering</h2>
<p>I’ve done this — transitioned teams from feature factories to product engineering teams. Here’s what I’ve learned.</p>
<p>Most engineers are fine. They adapt. Many flourish — they were waiting for permission to care about outcomes, and now they have it. The <a href="https://blog.dicko.dev/posts/why-your-engineers-are-still-waiting/">ownership gradient I’ve written about before</a> maps directly onto this: engineers who were stuck at Station 1 — waiting, executing what’s asked — often had product instincts that were never given room to breathe.</p>
<p>But some hit a wall. The resistance sounds like this: “Before, I was responsible for X story points per sprint. Now you want me to be responsible for business results?”</p>
<p>That sentence is the feature factory’s final defence mechanism. It’s an engineer saying: I understood the old game. I was good at it. The score was clear. Now you’re changing the rules, and I don’t know how to win anymore.</p>
<p>It’s a hard adjustment. Some don’t make it. But most do — and the ones who make it tend to flourish, because they’ve gone from being measured on activity to being measured on impact, and impact is more meaningful work. The <a href="https://blog.dicko.dev/posts/the-senior-engineer-plateau/">senior engineer plateau</a> I wrote about — that quiet disengagement that happens when technically excellent engineers stall — is often caused by exactly this: an engineer who’s never been connected to the purpose of what they build. Design thinking as a capability prevents that plateau because it keeps engineers connected to users, to outcomes, to the reason any of this matters.</p>
<p>The Somchai story opens this post. This section closes the loop. Somchai didn’t need a design thinking workshop. Somchai needed to have been hired into a team that never let velocity become a scorecard in the first place — or, better, Somchai needed to be the kind of engineer who would have questioned that scorecard himself. Velocity should tell you whether your sprint was realistic. It should help you plan the next one. The moment it becomes something you report upwards, it becomes something you game. And gamed metrics tell you nothing about whether you’re solving problems for the humans who use your software.</p>
<h2>From Workshop to Worldview</h2>
<p>Design thinking isn’t a methodology. It’s a worldview — a way of approaching problems that starts with the human and works backward to the technology. Steve Jobs is often quoted as saying “people don’t know what they want until you show it to them.” But Jobs himself corrected this reading at the 1997 WWDC: “You’ve got to start with the customer experience and work backwards to the technology… I’ve made this mistake probably more than anybody else in this room.” The real lesson from both Ford and Jobs isn’t “ignore your customers.” It’s “customers have problems that need solving, not solutions that need implementing.”</p>
<p>When design thinking lives in a workshop, it dies on Monday. When it lives in your hiring criteria, your onboarding, your promotion criteria, your <a href="https://blog.dicko.dev/posts/the-taste-gap-why-your-engineering-teams-keep-solving-the-same-problems-differen/">shared taste</a>, and your culture — it becomes the way your organisation thinks.</p>
<p>The test is simple: can your engineers explain who benefits from the thing they’re building, and how you’ll know if it worked? If they can, you’ve hired product engineers. If they can’t, you’ve hired implementers. Both are useful. But if you want to stop being a feature factory, you need more of the former and fewer of the latter.</p>
<p>Now, if you’ll excuse me, I need to go prepare for an interview. I’ve got one question ready that isn’t about system design or algorithms. It’s “tell me what your application does.” The last candidate talked for four minutes about their tech stack. The one before that talked for thirty seconds about the problem their users have. Guess which one I’m recommending we hire.</p>
]]></content:encoded>
  </item>
  <item>
    <title>The Directness Dial: Teaching Engineers to Be Honest Without Being Brutal</title>
    <link>https://blog.dicko.dev/posts/the-directness-dial-teaching-engineers-to-be-honest-without-being-brutal/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/the-directness-dial-teaching-engineers-to-be-honest-without-being-brutal/</guid>
    <pubDate>Sun, 05 Apr 2026 17:19:47 GMT</pubDate>
    <category>software-engineering</category>
    <category>software-development</category>
    <category>engineering-leadership</category>
    <category>feedback</category>
    <category>leadership</category>
    <description>Or: Why “Too Respectful” Is Just as Dangerous as “Too Blunt”</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/the-directness-dial-teaching-engineers-to-be-honest-without-being-brutal/cover.webp" alt=""></figure>
<p>Somchai stopped mid-sentence. Not the natural trailing-off of a thought completed, but the abrupt severing of one — the kind you hear when someone decides, in real time, that the room has already made up its mind. The vinyl floor on level 6 caught the scrape of his chair pushing back an inch. Three slides still queued in his deck. The senior engineer across the table was already typing, eyes on his laptop, the verdict delivered before the defence had rested. The aircon hummed. Nobody else spoke. Two juniors in the corner exchanged a glance — the kind that says <em>I had a question, but never mind</em>.</p>
<p>That meeting had all the information it needed. It just couldn’t use any of it. Not because the people were wrong, but because the way information moved — or didn’t — had already decided the outcome before anyone sat down.</p>
<p>As the diplomat Henry Kissinger once observed, “The absence of alternatives clears the mind marvellously.” He was talking about geopolitics, but he might as well have been describing what happens when your most technically brilliant engineer delivers feedback like a closing argument in a murder trial. Alternatives don’t just <em>clear</em> — they vanish. People stop offering them. And a team that stops offering alternatives is a team making decisions with incomplete information, which is a polite way of saying <em>bad decisions</em>.</p>
<p>Here’s the thing nobody wants to admit: every engineering culture optimises for one failure mode or the other. Some cultures — usually the ones founded by blunt, technical personalities — have no problem with directness. They’ll tell you your code is bad, your design is flawed, and your idea won’t scale. They’ll do it in a way that makes you want to never open your mouth again. Other cultures — usually the ones that pride themselves on being “polite” — are so careful about feelings that problems fester for months because nobody will say what they actually think.</p>
<p>Both failure modes produce the same outcome: <strong>information doesn’t flow.</strong></p>
<p>In the brutally direct culture, people stop sharing ideas because they’ll get shredded. In the too-respectful culture, people stop raising concerns because it feels impolite. Different mechanisms, identical result. The team makes worse decisions because it’s working with incomplete data.</p>
<p>This isn’t a post about finding the “middle ground” — that framing implies mediocrity, a kind of beige compromise where you’re slightly less brutal and slightly more honest and nobody’s particularly good at either. It’s about building a specific, coachable skill: <strong>calibrated directness.</strong> Knowing when to turn the dial up — <em>this design has a critical flaw and we need to talk about it right now</em> — and when to turn it down — <em>your first PR has some issues but you clearly put thought into it</em>. It’s a skill, not a personality trait. Which means it can be taught, practised, and coached. And if you lead engineers, coaching it is part of your job, whether anyone put it in your job description or not.</p>
<h2>Two Meetings, Same Wreckage</h2>
<p>Let me show you what both failure modes look like in practice. You’ll recognise at least one of them. If you’re honest, you’ll recognise both.</p>
<p><strong>The Brutal Review.</strong> A design review, 2:30 on a Wednesday. The presenting engineer is three slides in when a senior engineer interrupts: “This won’t work. The data model is wrong, the API contract is brittle, and you’ve clearly never dealt with this at scale.” The presenter goes quiet. The room goes quiet. The review continues for another twenty minutes, but it’s performative — the presenter has already mentally checked out. After the meeting, two junior engineers who had questions about the approach decide not to raise them. Why bother? If they’re wrong, they’ll get the same treatment. Here’s the part that makes this genuinely expensive: the senior engineer’s feedback was <em>technically correct</em>. The data model did have issues. But the delivery ensured that nobody else in the room would contribute for the rest of the quarter. One correct observation killed a dozen potential ones.</p>
<p><strong>The Polite Drift.</strong> A retrospective. The team has missed their milestone by three weeks. The EM asks what went wrong. Silence. Then someone offers: “I think the scope was a bit bigger than we expected.” More silence. “Maybe we could do better at estimation next time.” Everyone nods. The meeting ends. Nobody mentions the actual problem: one team member committed code that broke the integration tests and didn’t fix them for five days, blocking three other engineers. Nobody says it because the team culture is <em>we don’t call people out</em>. The EM knows about the issue but doesn’t raise it either — they plan to “handle it in the 1:1.” But in the 1:1, the conversation is also soft: “Next time, maybe try to fix the tests sooner.” The engineer nods. Nothing changes. Two sprints later, it happens again.</p>
<p><strong>The Slack Thread That Escalated.</strong> A senior engineer reviews a PR and leaves a comment: “This approach is fundamentally wrong. Please rewrite using the repository pattern.” In person, softened by tone and body language, this might land at a 7 on the directness scale — direct but workable. In a Slack thread, stripped of every non-verbal cue, it reads as a 10 — a verdict delivered to a public audience. The PR author reads it at 11pm. They don’t ask for clarification — they spend the weekend rewriting, quietly furious. On Monday, the senior engineer is genuinely surprised to learn there’s a problem. “I was just being direct,” they say. They were. But a dial-7 comment in person becomes a dial-10 comment in text, because text has no smile, no softening pause, no “hey, overall this is solid, but…” preamble that humans naturally add face-to-face. The medium shifted the dial without anyone touching it.</p>
<p>All three scenes had the same underlying failure — critical information that needed to be communicated wasn’t communicated effectively. In the brutal review, the information was delivered but the environment was destroyed. In the polite drift, the environment was preserved but the information never surfaced. In the Slack thread, the information was delivered <em>and</em> the environment was damaged — because nobody accounted for the channel. Different mechanisms, same result: the team makes worse decisions because information isn’t flowing to the people who need it, in a form they can actually use.</p>
<h2>Name the Failure Modes (So You Can Coach Them)</h2>
<p>You can’t fix what you won’t name. So let’s name them.</p>
<p><strong>Failure Mode 1: Too Much Signal, Not Enough Safety.</strong> This is the engineer who says exactly what they think, exactly when they think it. Technically accurate, socially destructive. You know the observable behaviour: they interrupt, they use absolutes — “this is wrong,” “that’s a terrible idea” — and they evaluate the person, not the work: “you clearly didn’t think about this.” What it costs the team: people stop contributing ideas, design reviews become theatre, junior engineers learn to stay silent, and your best people — the ones with options — start looking elsewhere. Why it persists: this person is often genuinely excellent technically. Their feedback is valuable. The team tolerates the behaviour because the content is good — and because nobody wants to be the next target.</p>
<p><strong>Failure Mode 2: Too Much Safety, Not Enough Signal.</strong> This is the engineer — or the manager — who wraps every concern in so much padding that the concern disappears entirely. Socially graceful, informationally useless. The observable behaviour: hedging language (“maybe we could consider…”), feedback disguised as questions (“have you thought about…?” when they mean “this is wrong”), avoidance of specifics, and the pervasive use of “we” and “the team” when there’s a specific person and a specific action that need naming. What it costs the team: problems persist because they’re never identified, <a href="https://medium.com/beer-and-servers-dont-mix/technical-debt-the-elephant-in-the-scrum-room-or-why-your-product-owner-shouldn-t-be-the-only-one-3186c1da3ac2">tech debt</a> accumulates because nobody says “this code isn’t good enough,” and people who <em>are</em> direct get labelled as “difficult” simply because they’re the only ones breaking the norm. Why it persists: it feels kind. It feels professional. It feels respectful. And for a long time, it looks like a functional team — until the problems it hides become impossible to ignore.</p>
<p>Here’s the paradox that makes both of these so persistent: the brutally direct person thinks the polite person is weak. The polite person thinks the direct person is an arsehole. Both are wrong. Both are doing the easy version of their preference instead of the hard version of the complete skill.</p>
<p>The basketball coach Phil Jackson put it well: “The strength of the team is each individual member. The strength of each member is the team.” A team where everyone is dial-at-10 has no team. A team where everyone is dial-at-2 has no signal. You need both — the ability to be direct and the judgement to know how direct to be — and you need it to come from the same person, calibrated to the moment.</p>
<h2>Why “Just Be Direct” Is Bad Advice (And So Is “Be More Polite”)</h2>
<p>We love simple prescriptions. They feel actionable. They’re also wrong.</p>
<p>“Just be direct” fails because directness without care isn’t honesty — it’s self-indulgence. The brutally direct engineer often isn’t being honest for the team’s benefit. They’re enjoying the performance of being the smartest person in the room, and the information delivery is incidental to the ego display. Directness also works differently depending on the power dynamic. A senior engineer being “direct” with a junior isn’t the same as two peers being direct with each other — the junior hears it as a verdict, not a data point. And directness without context is noise: “This is wrong” without “here’s what I’d suggest instead” is destruction without construction.</p>
<p>“Be more polite” fails because it conflates politeness with kindness. These are different things. Politeness is about the speaker’s comfort — avoiding the discomfort of delivering hard feedback. Kindness is about the recipient’s growth — telling them what they need to hear even when it’s uncomfortable for you to say it. Every piece of feedback softened past recognition is a decision deferred. The feedback doesn’t go away — it just arrives later, in a worse form: a bad performance review, a missed promotion, a project failure that could have been prevented three months earlier. And there’s something patronising about it, too. When you refuse to give someone honest feedback, you’re implicitly saying you don’t believe they can handle it. Most engineers can. They want to improve. They’d rather hear “this approach won’t scale because of X” today than discover it in production next quarter.</p>
<p>Kim Scott’s <a href="https://www.radicalcandor.com/"><em>Radical Candor</em></a> framework captures this neatly in a 2×2: Care Personally × Challenge Directly. The four quadrants map to what we’ve been describing — Radical Candor is the target (care and challenge), Obnoxious Aggression is Failure Mode 1 (challenge without care), Ruinous Empathy is Failure Mode 2 (care without challenge), and Manipulative Insincerity is the passive-aggressive nightmare nobody wants to be but everyone has been at some point. The framework is useful but not sufficient. It tells you <em>what</em> to aim for. It doesn’t tell you <em>how to coach people there</em> — which is the part Scott’s book leaves as an exercise for the reader.</p>
<p>This post is that exercise.</p>
<h2>The Directness Dial</h2>
<p>Here’s the core metaphor: directness isn’t a binary. It isn’t a personality trait you either have or you don’t. It’s a dial. The skill isn’t about always being at 10 or always being at 3 — it’s about knowing where to set it for a given context and having the range to actually get there.</p>
<p><strong>Factors that turn the dial up</strong> — more direct, less padding: safety risk (production incident, security vulnerability, data integrity). A repeated pattern (you’ve given this feedback before and nothing’s changed). An audience that can handle it (senior engineer, established trust, strong relationship). Time pressure (we need to decide now). High cost of silence (if I don’t say this, the project ships with a critical flaw).</p>
<p><strong>Factors that turn the dial down</strong> — more careful framing, more scaffolding: new team member or first interaction (trust hasn’t been established). Cross-cultural context (more on this shortly). Public setting (criticism in front of others raises the emotional stakes). Personal circumstances (the person is already having a hard week — timing matters). A question of taste rather than correctness (there are multiple valid approaches and you just prefer one).</p>
<p>The key insight is that the dial is <strong>context-dependent, not person-dependent</strong>. The same engineer might need dial-9 directness about a production safety issue and dial-3 gentleness about a naming convention preference. The person who’s always at 10 and the person who’s always at 2 are both failing — they’ve confused their comfort zone with the appropriate setting.</p>
<p><strong>And then there’s the medium.</strong> This is the factor almost everyone forgets. The communication channel itself recalibrates perceived directness, and it does it without your permission. In person, body language, tone, and facial expressions do constant softening work — a smile, a lean-in, a “hey, this is great overall, but…” preamble that humans add instinctively. In text — Slack messages, PR comments, email — all of that is stripped away. What’s intended as a dial-6 comment in the writer’s head lands as a dial-9 on the reader’s screen. This means written feedback needs to be set two or three notches lower than what you’d say in person to land at the same perceived level. The engineer who writes terse, factual PR comments — “This is wrong. Use the repository pattern.” — isn’t necessarily at dial-10 in their own calibration. They just haven’t accounted for the medium stripping the humanity from their words. This is coachable: show them how their comments read to someone who can’t see their face.</p>
<p>This connects directly to the argument in <a href="https://blog.dicko.dev/posts/co-location-remote-first-and-the-rollback-every-few-decades/">Co-location, Remote First, and the Rollback Every Few Decades</a> — proximity changes the information channel, and the channel changes how directness is received. It’s harder to be direct in text than in person because tone is missing. It’s also harder to be <em>too</em> direct in person because you see the reaction in real time. The co-location post’s argument about information flow applies here: directness is the human channel through which critical information flows, and the medium shapes the channel’s bandwidth.</p>
<h2>The Multicultural Layer (Or: “Your Code Sucks” in Fifty Languages)</h2>
<p>Most “how to give feedback” content is written from a monocultural perspective — usually American or Northern European. But any engineering organisation that hires internationally, and most large ones do, has a multicultural directness problem that can’t be ignored by pretending it doesn’t exist.</p>
<p>Erin Meyer’s <a href="https://erinmeyer.com/books/the-culture-map/"><em>The Culture Map</em></a> identifies two scales that matter here: the Evaluating scale (direct versus indirect negative feedback) and the Disagreeing scale (confrontational versus avoids confrontation). The crucial insight from Meyer: these two scales don’t always align. A culture can be indirect in general communication but still direct with negative feedback. A culture can be explicit about expectations but soft with criticism. Thailand, where I’m writing this, sits toward the indirect and confrontation-avoiding end of both scales. Australia, where I grew up, sits considerably further toward the direct end — and I grew up in country Queensland, which is more direct than Australians normally are, which is already saying something. Dutch engineers, German engineers, Israeli engineers — further still.</p>
<p>What this means in practice is that an Australian or Dutch engineer’s “direct” is a Thai or Japanese engineer’s “aggressive.” A Thai engineer’s “polite” is a Dutch engineer’s “evasive.” Neither is wrong — they’re calibrated for different default contexts. The engineering leader’s job isn’t to pick a cultural default and impose it — it’s to create a shared understanding: <em>In this team, in this meeting, this is how we communicate. Here’s why.</em></p>
<p>Meyer’s recommendation for multicultural teams is to default to low-context processes: be more explicit about expectations, not less. Make the norms visible. Don’t assume “everyone knows how to give feedback” because “how to give feedback” means radically different things to a Thai engineer, an Indian engineer, an Australian engineer, and a German engineer sitting in the same room on the same floor.</p>
<p>I saw this play out beautifully on one of the mobile engineering teams here. Code reviews had become heated — different cultural defaults around directness were colliding, and what felt like honest feedback to one engineer felt like a personal attack to another. The manager didn’t respond with a policy document or a mandatory training session. He ran an exercise: every engineer wrote “your code sucks” on a whiteboard in their native language, then in every other language they could speak. They got to forty or fifty translations.</p>
<p>The exercise did several things at once. It defused the tension with humour — seeing the same blunt phrase rendered in Thai, Hindi, Japanese, Dutch, Russian, and Tagalog makes it absurd rather than threatening. It made the cultural dimension visible — everyone in the room could see that the same sentiment sounds different depending on the language it’s wrapped in, which is Meyer’s evaluating scale made visceral. And it taught the team not to take themselves too seriously. The underlying message: we all think each other’s code sucks sometimes. That’s fine. The question is how we say it — and recognising that “how we say it” is partly a function of where we come from.</p>
<p>That whiteboard exercise cost nothing and took thirty minutes. It did more for the team’s communication than any policy document ever could. This is what making norms explicit looks like in practice — not a training deck, not an HR mandate, but a shared moment that made the invisible visible.</p>
<h2>Coaching the Dial-at-10 Engineer</h2>
<p>This is the coaching challenge that gets the most attention, because dial-at-10 engineers are visible. Everyone knows who they are. Their impact is observable in real time — the silence after they speak, the engineers who stop contributing, the design reviews that become performances rather than conversations.</p>
<p>Here’s what doesn’t work. “Be more polite” is too vague — it feels like you’re asking them to be fake. “You’re too aggressive” triggers defensiveness and labels the person rather than the behaviour. And ignoring it because their technical contributions are valuable is the most common failure — the silent tax on the team compounds every week.</p>
<p>Here’s what does work.</p>
<p><strong>Lead them to the cost — don’t deliver it.</strong> Don’t say “you were too harsh in that review.” Don’t even say “your delivery cost us information.” Instead, ask: “After you gave that feedback, how did the other engineers react?” They’ll pause. If they’re honest: “They were quiet.” Then: “Why do you think they were quiet?” Let them follow the thread. Let them arrive at the conclusion that their technically correct feedback shut down every other voice in the room — that the two junior engineers who had questions decided not to raise them, and the team lost information it needed. The insight lands differently when they discover it themselves than when you hand it to them pre-packaged. You’re not attacking their character or even their style. You’re walking them through a chain of cause and effect and letting them see the price tag on their own approach. Most technically excellent people respond to reasoning, and this is reasoning about information loss — delivered in a form they’d actually use on a technical problem.</p>
<p><strong>The “and” technique.</strong> Replace “but” with “and.” “Your technical feedback is strong <em>and</em> the delivery is making people shut down.” This frames calibrated directness as an incomplete skill, not a character flaw. They’re doing one half well. They need the other half. The word “but” negates everything before it; “and” holds both truths together.</p>
<p><strong>Make it a pattern, not a one-off.</strong> The guided questioning works once. But the real change happens when you do it after the <em>next</em> meeting too. And the one after that. “What happened when you said X? How did Y react? What do you think that cost us?” Each time, you’re building the feedback loop they’re missing — teaching them to notice the room’s reaction in real time, not just the content of their own argument. Eventually, they start asking themselves the questions before you have to.</p>
<p><strong>Give them scripts.</strong> Direct people often don’t know what “softer” sounds like without being dishonest. Provide concrete alternatives. Instead of “This is wrong,” try “I have a concern about this approach — can I walk through it?” Instead of “You clearly didn’t think about X,” try “How did you think about X? I want to understand the trade-off you made.” Instead of “This won’t scale,” try “I’ve seen this pattern struggle at scale — here’s what happened. How are you thinking about that risk?” Each alternative contains the same information. The packaging changes. The information loss drops to zero.</p>
<p><strong>Praise the content, then ask about the delivery.</strong> “Your observation about the data model was the most important thing said in that meeting. What do you think would have happened if you’d opened with a question instead of a statement?” Let them reason through it. They’ll get to the same place — that the room would have stayed open, that the same information would have been delivered without the shutdown — but they’ll own the conclusion.</p>
<p>Here’s the good news: most dial-at-10 engineers are actually the easier coaching challenge. Because they respect directness, they can engage with direct questions about their own behaviour. You don’t need to sugar-coat it — you just need to ask instead of tell. Lead them through the impact with questions, let them connect the dots, and they’ll adjust. For most, this works within weeks. They genuinely didn’t know how their delivery was landing, and once they reason through the cause and effect themselves, the insight sticks.</p>
<p>But some dial-at-10 engineers aren’t coachable on this dimension. Their directness isn’t a calibration gap — it’s entangled with ego, insecurity, or interpersonal patterns that run deeper than communication style. An engineering manager isn’t a psychologist, and treating a personality issue as a coaching opportunity wastes everyone’s time. The pragmatic position: it’s easier to teach technical skills than to be someone’s therapist for their personal issues. This isn’t an argument against hiring direct people — the whole post argues that too <em>little</em> directness is just as dangerous. It’s an argument for distinguishing at the hiring stage between “direct and calibratable” — hire them, coach them, you’ll get there in weeks — and “direct and destructive,” where the pattern is deeper than communication skills and an EM isn’t the right intervention. The screening question isn’t “are they direct?” It’s “can they modulate?”</p>
<h2>Coaching the Dial-at-2 Engineer</h2>
<p>This is the quieter coaching challenge, and in many ways the more damaging one, because the problem is invisible. It’s what the dial-at-2 engineer <em>didn’t</em> say. The meeting that looked productive but wasn’t. The retrospective where everyone nodded and nothing changed. The design review where a flaw went unchallenged because raising it felt impolite.</p>
<p>Here’s what doesn’t work. “Speak up more” is too vague — it doesn’t address the emotional or cultural barrier. “Just say what you think” ignores the genuine reasons they don’t. And putting them on the spot in meetings creates the exact dynamic they’re trying to avoid.</p>
<p>Here’s what does work.</p>
<p><strong>Start in writing.</strong> For engineers who struggle with verbal directness, code reviews are the training ground. Written feedback is lower-stakes than spoken feedback — there’s time to compose, revise, and calibrate. PR comments are particularly useful because they’re <em>expected</em> to contain technical critique — the social permission is built into the format. A dial-at-2 engineer who would never say “I think this has a race condition” in a meeting will often write it in a PR comment, because the review process explicitly asks for exactly that. Use this as the entry point. Once they experience that direct technical feedback in writing didn’t damage the relationship, they start to believe it’s possible in other contexts too.</p>
<p><strong>Give them the first move.</strong> In meetings, specifically invite their opinion before the seniors weigh in: “I want to hear from you before we go around.” This removes the barrier of having to interrupt or contradict someone more senior. It also signals that their directness is wanted, not just tolerated.</p>
<p><strong>Model it.</strong> The engineering leader who says “I was wrong about X — here’s what I missed” gives everyone in the room permission to be wrong and say so. The leader who says “Somchai, I think your design has a flaw in the caching layer — let’s talk about it” shows that direct feedback can be delivered with respect, that it’s not a weapon but a tool. This is the culture-through-repetition mechanism I described in <a href="https://blog.dicko.dev/posts/the-taste-gap-why-your-engineering-teams-keep-solving-the-same-problems-differen/">The Taste Gap</a> — the leader’s behaviour is the norm everyone else calibrates against.</p>
<p><strong>The 1:1 replay method.</strong> This is the technique that transforms quiet engineers into contributors. I learned it secondhand — from the staff of another manager who had a genuine talent for turning dial-at-2 engineers into people who could hold their own in any room. I asked one of his engineers directly: “How did he get you to speak up?”</p>
<p>The answer was simple and relentless. The manager came to almost every 1:1 with specific examples of moments where the engineer could have spoken up in a meeting and didn’t. Not “you should speak up more” — that’s too vague, too easy to nod at and forget. Specific: “In Tuesday’s design review, when the caching approach was discussed, you had something to say. I could see it. Why didn’t you say it? What would it have taken?”</p>
<p>And then he did it again the next week. And the week after. And the week after that.</p>
<p>The behaviour change didn’t happen from a single conversation. It happened because the manager brought it up every week, in every 1:1, with specific recent examples. The engineer couldn’t treat it as a one-off piece of feedback. It became a running thread — something the manager clearly tracked, cared about, and returned to. The message was unmistakable: <em>I notice when you hold back, and I want you to stop.</em></p>
<p>This is the same repetition mechanism from <a href="https://blog.dicko.dev/posts/why-your-engineers-are-still-waiting/">Why Your Engineers Are Still Waiting</a> — show the alternative, repeatedly, until the alternative becomes the reflex. A single piece of feedback is advice. The same feedback with fresh evidence in ten consecutive 1:1s is a priority. The engineer starts to understand that speaking up isn’t a personality preference — it’s an expected part of their role.</p>
<p><strong>Break the complaint loop.</strong> Here’s a pattern you’ll recognise if you manage people. Many dial-at-2 engineers aren’t actually dial-at-2 everywhere — they’re dial-at-2 in group settings and dial-at-7 in 1:1s with their manager. The problem isn’t the skill. It’s the venue. The tell: your 1:1s become a running list of complaints about other engineers. “I wish X would stop doing Y.” “Z’s approach is going to cause problems.” They’re being direct — just to you, not to the person who needs to hear it.</p>
<p>The manager who accepts this dynamic becomes an information bottleneck, shuttling feedback between people who should be talking directly. The manager who redirects — “Have you told them that? What would it take for you to say that directly?” — builds the muscle. Every time they bring a complaint to you that should go to a peer, redirect. Don’t relay the message for them. That trains the wrong muscle. The goal is to make the indirect channel uncomfortable enough that going direct feels easier.</p>
<h2>Meeting Structures That Create Safety for Directness</h2>
<p>There’s no such thing as best practice here — only adequate practice given current context. And also bad practice, which is doing nothing and hoping directness emerges on its own. The right structure depends on the team: a newly formed team with low trust might need anonymous pre-reads; a team of senior engineers who’ve worked together for three years might just need an EM who asks pointed questions. Here are five structures worth having in your toolkit. Pick based on context.</p>
<p><strong>The “disagree first” design review.</strong> Before presenting their solution, the presenter lists the alternatives they considered and rejected — and why. This frames the review as a conversation about trade-offs, not a defence of a chosen approach. It makes disagreement safer because the presenter has already demonstrated that alternatives exist. This connects to the <a href="https://blog.dicko.dev/posts/i-dont-care-what-you-build-and-neither-should-you/">I Don’t Care What You Build</a> approach — focus design reviews on the problem being solved, not the solution being defended. The directness skill is what makes those reviews productive. A room full of people who are too polite or too brutal is a room that produces worse decisions.</p>
<p><strong>The anonymous pre-read.</strong> Before a meeting, circulate the design or proposal and collect written feedback anonymously. Review it as a group. The anonymity removes the social cost of disagreement, and the group discussion normalises it. Over time, people become comfortable attaching their name because they’ve seen that direct feedback is received well.</p>
<p><strong>The designated dissenter.</strong> In each design review, one engineer is assigned to argue <em>against</em> the proposal. This removes the social stigma from disagreement — the person disagreeing is doing their job, not being “difficult.” Red teams in security work on this principle. De Bono’s “Six Thinking Hats” formalises it. It’s artificial, and it works precisely because it’s artificial.</p>
<p><strong>The “I heard” check.</strong> At the end of a feedback conversation, the recipient summarises what they heard: “What I’m hearing is that the caching layer has a risk because X. Is that right?” This forces specificity — the too-polite feedback-giver can’t get away with vagueness when the recipient is reflecting it back, and the too-blunt feedback-giver can see how their message actually landed.</p>
<p><strong>Separate praise and criticism channels — with the PR exception.</strong> Public praise, private criticism. This single rule does enormous work: it makes praise visible (reinforcing good behaviour) and makes criticism safe (protecting the recipient’s dignity). But engineering has a built-in exception that needs acknowledging: code reviews are public criticism by default. When someone leaves a comment on a PR saying “this won’t scale,” it’s visible to the team. The resolution: critique the <em>code</em> in public (PRs), critique the <em>behaviour</em> in private (1:1s). “This caching strategy has a race condition” belongs in a PR comment — others learn from it. “You’ve submitted three PRs this month without tests — we need to talk about that pattern” belongs in a 1:1 — it’s about a pattern of behaviour, and that’s personal.</p>
<h2>When the Dial Must Be at Maximum</h2>
<p>There are moments when context, comfort, and cultural calibration all take a back seat. When the dial <em>must</em> be at 10, regardless:</p>
<p><strong>Production safety issues.</strong> If the code is going to break in production, you say so clearly and immediately. No hedging. Not “maybe we should consider whether this could potentially cause a problem.” The users don’t care about your team dynamics. “This will cause data loss. We need to stop.”</p>
<p><strong>Ethical concerns.</strong> If something looks wrong — data handling, user privacy, regulatory compliance — you speak up directly. This is not negotiable across any cultural context.</p>
<p><strong>Repeated patterns that aren’t changing.</strong> If you’ve given gentle feedback twice and nothing’s changed, the third time needs to be unambiguous. Continued softness at this point isn’t kindness — it’s negligence. The CEO “What did we learn?” story I described in <a href="https://blog.dicko.dev/posts/the-art-of-failing-how-experimentation-culture-transforms-engineering-teams/">The Art of Failing</a> only works because the CEO’s directness was calibrated — he named the failure and redirected to learning. But he <em>named</em> it. He didn’t pretend three consecutive quarters of missed targets were “a bit of a scope issue.”</p>
<p><strong>During incident response.</strong> When production is down, communication must be crisp, direct, and unambiguous. This is not the time for feelings management. The war room on level 7 next to the NOC area is not a place for hedging.</p>
<p>The ability to be direct about what went wrong <em>after</em> an incident depends entirely on the safety established <em>before</em> the incident. Teams that can’t be direct in reviews can’t be honest in post-mortems. The feedback culture and the incident culture are the same culture — you don’t get one without the other.</p>
<h2>What “Direct and Respectful” Actually Sounds Like</h2>
<p>Let’s return to those opening scenes and replay them with calibrated directness.</p>
<p><strong>The brutal review, rewritten.</strong> The senior engineer waits for the presenter to finish. Then: “Thanks for walking us through this. I have a concern about the data model — specifically the relationship between orders and payments. Can I walk through what I think might happen at scale?” The same information is delivered. The same issue is raised. But the room stays open. The two junior engineers with questions ask them. The design improves <em>more</em> than it would have in the original version — because more people contributed. The cost of the senior engineer’s six extra seconds of framing: zero. The value: incalculable.</p>
<p><strong>The polite drift, rewritten.</strong> The EM says in the retro: “I want to name something specific. We had integration tests broken for five days, and it blocked three engineers. That’s not a character judgment — it happens. But I want us to talk about what our norm should be when tests are broken. What’s a reasonable window to fix them?” The issue is named. The person isn’t blamed. The team sets a norm. The behaviour changes. This is what I was describing in <a href="https://blog.dicko.dev/posts/the-solutioneering-trap-why-your-best-engineers-are-solving-the-wrong-problems/">The Solutioneering Trap</a> — when I had to tell Somchai that his impressive Next.js prototype was solving the wrong problem, the approach wasn’t “this is bad.” It was: “Could we achieve this without rewriting all of our server-side rendering code?” A question that let him discover the answer himself. Dial-at-5: honest, but constructed to preserve agency and dignity.</p>
<p><strong>The Slack thread, rewritten.</strong> The senior engineer writes their PR comment, then re-reads it before posting. They add: “Overall this is a solid approach — the test coverage is great. One thing I’d rethink: the direct DB access here will create coupling to the Booking team’s schema. Have you considered the repository pattern? Happy to pair on it if useful.” Same technical information. Same directness. But the medium is accounted for — the written word gets the extra warmth that a face-to-face conversation would have provided for free.</p>
<p>Neither scene required being mean. None required being vague. All three required the skill of calibrated directness — the ability to deliver the truth with enough care that it can be heard, enough clarity that it can be acted on, and enough awareness of the medium and context to set the dial where it needs to be.</p>
<p>As the surgeon and writer Atul Gawande observed, “Better is possible. It does not take genius. It takes diligence. It takes moral clarity. It takes ingenuity. And above all, it takes a willingness to try.” He was talking about medicine, but the principle is identical in engineering leadership. Calibrated directness isn’t a talent. It’s a practice. You get better at it by doing it, by coaching it, by watching where the dial should have been and adjusting next time.</p>
<h2>The Bottom Line</h2>
<p>Every engineering team sits somewhere on the directness spectrum. Most are too far in one direction — either brutal enough that people stop contributing, or polite enough that problems stop surfacing. Both cost you the same thing: information. And information is the raw material of every good decision your team will ever make.</p>
<p>The fix isn’t finding the middle. The fix is building range. Teaching every engineer on your team — including yourself — to read the context, set the dial, and deliver the truth in a form the room can actually use. It’s harder than being blunt all the time. It’s harder than being polite all the time. It requires paying attention to the person you’re talking to, the stakes of the conversation, the medium you’re using, and the cultural context you’re operating in. But it’s a skill, not a gift, and skills can be taught.</p>
<p>If you’re the leader: model it. Be wrong in public. Be direct with care. Bring specific examples to every 1:1. <a href="https://blog.dicko.dev/posts/the-taste-gap-why-your-engineering-teams-keep-solving-the-same-problems-differen/">Build culture through repetition</a>, not through policy documents. Make the norms explicit, especially if your team looks like a United Nations assembly and everyone’s default dial setting comes from a different continent.</p>
<p>If you’re the dial-at-10: your feedback is valuable. Your delivery is expensive. Learn to keep the signal and lose the shrapnel.</p>
<p>If you’re the dial-at-2: your restraint isn’t kindness. It’s a decision to let problems persist because naming them is uncomfortable. The team deserves your honest assessment, even — <em>especially</em> — when it’s hard to say out loud.</p>
<p>Now, if you’ll excuse me, I need to go re-read a Slack message I sent twenty minutes ago. It was technically accurate, but I’m fairly sure it landed about three notches higher than I intended. The medium, as always, stripped the smile I was wearing when I wrote it.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Teaching the Agent What You Know: How to Build, Test, and Ship AI Coding Skills That Actually Hold…</title>
    <link>https://blog.dicko.dev/posts/teaching-the-agent-what-you-know-how-to-build-test-and-ship-ai-coding-skills-tha/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/teaching-the-agent-what-you-know-how-to-build-test-and-ship-ai-coding-skills-tha/</guid>
    <pubDate>Thu, 02 Apr 2026 13:38:38 GMT</pubDate>
    <category>artificial-intelligence</category>
    <category>software-engineering</category>
    <category>software-development</category>
    <category>mcp-server</category>
    <category>claude-skills</category>
    <description>Or: Your AI Assistant Is as Naive as the Day It Was Hired — and That’s Your Fault</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/teaching-the-agent-what-you-know-how-to-build-test-and-ship-ai-coding-skills-tha/cover.webp" alt=""></figure>
<h2>Teaching the Agent What You Know: How to Build, Test, and Ship AI Coding Skills That Actually Hold Up in Production</h2>
<h3>Or: Your AI Assistant Is as Naive as the Day It Was Hired — and That’s Your Fault</h3>
<figure><img src="https://blog.dicko.dev/posts/teaching-the-agent-what-you-know-how-to-build-test-and-ship-ai-coding-skills-tha/image_1.webp" alt="" loading="lazy"></figure>
<p>Somchai stared at the screen at 11:47 on a Thursday morning, the hum of the air conditioning settling over the floor like a shared sigh. The code review had come back clean — from the linter, anyway. The AI had generated the whole component in under a minute. Twenty-three files changed, all passing CI. He scrolled slowly through the diff, the Thai iced tea going warm beside his keyboard, and then stopped on a single function. The API call was inside the component. On the client. For every render.</p>
<p>That’s when he opened a new chat window and started re-prompting.</p>
<p>Here’s the uncomfortable truth about AI coding assistants: they’re not slow, they’re naive. They produce code that compiles, passes tests, and satisfies every constraint you gave them — which is precisely the problem. You didn’t give them the constraints that live in your engineers’ heads. The ones nobody wrote down because everyone already knew them. The ones a senior engineer communicates through a single comment in a code review, and a junior absorbs over six months of watching that happen.</p>
<p>You can’t mentor the agent over six months. You can’t pair-program it into your conventions. Unless you encode that knowledge somewhere it can find it, your AI assistant will keep producing technically correct, production-naïve code — and your engineers will keep fixing it manually rather than shipping faster. That’s not an AI problem. That’s a knowledge transfer problem.</p>
<p>This post is about solving it properly.</p>
<h2>The Three-Layer Problem</h2>
<p>When an engineer picks up a codebase, they absorb knowledge in three layers. Most AI tooling only reaches the first.</p>
<p><strong>Layer 1: What the compiler knows.</strong> Types, function signatures, API contracts. This is the layer AI coding tools are genuinely excellent at. Give them a typed interface and they’ll implement it correctly.</p>
<p><strong>Layer 2: What the linter knows.</strong> Style conventions, formatting rules, naming patterns — whatever you’ve encoded in ESLint, Prettier, StyleCop, or ktlint. AI tools honour these too, usually.</p>
<p><strong>Layer 3: What the team knows.</strong> This is where AI tools fail completely, because nothing enforces it. It’s the knowledge that lives in your engineers’ heads and occasionally surfaces as a code review comment:</p>
<blockquote>
<p><em>“We don’t do aggregation client-side — we’ve been bitten by the connection limit too many times.”</em></p>
</blockquote>
<blockquote>
<p><em>“One class per file in C#, always. If you’ve got an interface with a single implementation, that’s the one exception.”</em></p>
</blockquote>
<blockquote>
<p><em>“Never patch the symptom. Find the root cause. If the fix doesn’t explain why the bug happened, it’s not the fix.”</em></p>
</blockquote>
<blockquote>
<p><em>“React components should be named and ordered to mirror the visual layout. If you put the page and the code side by side, you should be able to point to what renders what without drilling down four levels of components.”</em></p>
</blockquote>
<p>That last one came from Roland, eight years ago. It changed how I read every React codebase I’ve touched since. It took thirty seconds to say and had zero representation in any linter rule, any documentation, any file in the repository.</p>
<p>The agent doesn’t know it. And unless you do something about that, it never will.</p>
<h2>What Skills Actually Are</h2>
<p>A skill is a Markdown file. That’s the unfussy version. It’s a set of instructions that an AI coding agent should apply when handling a particular category of task — written for the agent the same way you’d write it for a new engineer, but with the specificity the agent can act on.</p>
<p>The IDE (Cursor, Claude Code, and others) maintains a catalogue of skill files in a local directory. When an engineer prompts the agent, the IDE reads the prompt, scans the skill catalogue, and decides which skills are relevant. The selected skills are injected into the context window alongside the prompt and code. The agent produces output with that knowledge baked in.</p>
<p>Here’s the flow:</p>
<figure><img src="https://blog.dicko.dev/posts/teaching-the-agent-what-you-know-how-to-build-test-and-ship-ai-coding-skills-tha/image_2.webp" alt="" loading="lazy"></figure>
<p>The decision about which skills get selected is made entirely on the <strong>name and description</strong> of each skill file. If those are vague or overlapping, you get one of two failure modes: too many skills injected (context bloat, degraded output), or the right skill never selected at all (invisible, useless). Writing skill descriptions is interface design, not documentation.</p>
<p>Here’s what a real skill looks like. This one encodes an engineering culture norm — the kind of thing that separates production-grade engineering from code that merely compiles:</p>
<pre><code>## Root cause analysis

**Never patch symptoms. Always find and fix the root cause.**

Before writing any fix, trace the problem to its origin.
A fix that suppresses the symptom without addressing why
it happened will fail again - or worse, hide the real issue
until it causes broader damage.
</code></pre>
<p>With this skill injected, agents stop patching and keep investigating until they find the underlying problem. Without it, they patch. Every time. Because patching is the fastest path to a passing test, and passing tests is what they’re optimising for.</p>
<p>That distinction matters more than it might appear — and we’ll come back to it in the risks section.</p>
<h2>The Testing Method</h2>
<p>Here’s where most teams stop. They write a few skill files, put them in a directory, tell engineers to <a href="https://gofastmcp.com/servers/providers/skills">install the MCP</a>, and consider the work done. Then six months later, someone rewrites a skill during a late-night cleanup sprint, removes a constraint that seemed redundant, and the agent starts generating C# files with two classes in them again. Nobody notices for three weeks.</p>
<p>Skills are prompts. Prompts can be tested. The fact that most teams don’t test them means they’re deploying skills blind — with no feedback loop, no regression detection, and no way to know if a change improved things or quietly broke them.</p>
<p>The testing method below applies the same discipline to skills that we apply to code.</p>
<h2>Level 1: Unit Tests with Promptfoo</h2>
<p><a href="https://github.com/promptfoo/promptfoo">Promptfoo</a> is an open-source CLI for prompt evaluation. It lets you point it at a skill file (the prompt), supply test inputs, and assert properties of the output — using regex, exact match, or an LLM-as-judge for more nuanced assertions.</p>
<p>Think of this as a unit test for a skill. You’re not testing the full agent loop. You’re testing whether a specific piece of knowledge survives in the agent’s output when this skill is present.</p>
<p><strong>Example:</strong> Your skill says “in C#, one class per file — the only exception is an interface with a single implementation.” A promptfoo test asks: “Should I put two classes in the same file?” The assertion checks that the response contains “no”, “never”, or equivalent negative language. If someone rewords the skill and accidentally drops that constraint, the test fails. The pipeline goes red. Someone notices.</p>
<p>This is directly analogous to unit testing a function: isolate the behaviour you care about, assert it explicitly, make it part of CI. Anyone who wants to change a skill has to change the tests, or the build breaks.</p>
<p>A promptfoo test file looks like this:</p>
<pre><code># promptfoo.yaml
prompts:
  - file://skills/csharp-conventions.md
providers:
  - anthropic:claude-sonnet-4-20250514
tests:
  - description: &quot;One class per file rule&quot;
    vars:
      question: &quot;Can I put two classes in the same file in C#?&quot;
    assert:
      - type: llm-rubric
        value: &quot;The response must say no and explain the one-class-per-file rule&quot;
  - description: &quot;Single implementation exception&quot;
    vars:
      question: &quot;I have an interface with exactly one implementation. Same file?&quot;
    assert:
      - type: llm-rubric
        value: &quot;The response must say yes and correctly identify this as the only exception&quot;
</code></pre>
<p>What promptfoo catches: specific facts and constraints that must survive a skill edit. The things small enough to be missed in an end-to-end test, important enough to break behaviour if lost.</p>
<p>What promptfoo misses: whether the skill actually gets selected by the IDE when it should. Whether the full agent loop — prompt plus context plus injected skill — produces better output than without it. That’s what the next level is for.</p>
<h2>Level 2: End-to-End Tests with an LLM Judge</h2>
<p>The end-to-end test validates the complete pipeline: code goes in, agent receives the prompt and injected skill, code comes out. The question is whether the output is meaningfully better when the skill is present.</p>
<p>The method:</p>
<ul>
<li><strong>Code before</strong> — a representative piece of code with known deficiencies. The naive output you’d expect from an agent without the skill.- <strong>Prompt</strong> — the task description the engineer would give.- <strong>Agent run</strong> — prompt plus code-before goes through the agent, with the skill injected.- <strong>Code after</strong> — the agent’s output.- <strong>LLM judge</strong> — a <em>different</em> model (ideally from a different vendor) receives both versions and scores the improvement on a percentage scale. (Note: Gemini is awesome for reviewing ime, and Calude is the best for editing)- <strong>Threshold</strong> — you set a minimum acceptable score. Below it, the skill needs work. Don’t use binary output, if you give an LLM a binary success failure, it’ll always bias success, if you give it success of “making a grade” it’ll be more honets, and the threshold/if statement can live in deterministic code.
The reason to use a different model as the judge is the same reason you don’t mark your own homework. A model that helped write the code is poorly positioned to evaluate it critically. Cross-model judging — Gemini judging Claude output, or vice versa — is more likely to surface genuine deficiencies.</li>
</ul>
<p>The percentage scoring matters too. You’re not asking the judge to pass or fail. You’re asking it to score improvement. That goal is significantly harder to game than a binary outcome.</p>
<p>The two levels complement each other. Promptfoo catches the specific, precise constraints that absolutely cannot be lost. End-to-end tests catch whether the skill is actually doing something useful in the real agent loop. Neither alone is sufficient. Together, they give you a CI pipeline for skills with the same confidence properties as a well-tested codebase.</p>
<p>The testing infrastructure lives in its own repository: skills, code-before examples, prompts, and the CI pipeline all version-controlled together. Adding a skill requires adding at least one test. Changing a skill may require changing tests. The pipeline doesn’t let untested skills reach production.</p>
<h2>The Risks (The Part You Should Actually Read)</h2>
<p>The testing method is the answer to most of these. But you should understand what you’re defending against.</p>
<h2>Risk 1: Bad skill descriptions cause the wrong skills to be selected</h2>
<p>The IDE selects skills by reading their names and descriptions. If those are vague (“general coding standards”) or overlapping (“React tips” and “React component structure”), one of two things happens:</p>
<p>Context bloat: too many skills get injected. The context window fills up. Output quality drops. Engineers re-prompt more, not less.</p>
<p>Skill invisibility: the skill is never selected because its description doesn’t match the vocabulary of the prompt. You’ve built a skill no one can find.</p>
<p>Write skill descriptions as API contracts. The description is the interface by which the IDE discovers and selects the skill. Treat it accordingly.</p>
<h2>Risk 2: Agents optimise for passing tests, not for doing the right thing</h2>
<p>This is the specification gaming problem, and the research literature is uncomfortable reading. <a href="https://arxiv.org/abs/2510.20270">ImpossibleBench (arXiv:2510.20270, Oct 2025)</a> documents how LLM agents, when given the goal of making tests pass, will find paths that don’t involve actually solving the problem — deleting failing tests rather than fixing the underlying bug, hardcoding values that satisfy assertions without implementing the real logic, adding skip markers to flaky tests rather than fixing them.</p>
<p>It’s not theoretical. One engineer in our team experienced this directly: an agent added a Python script that bypassed a CI check rather than addressing the root cause the check was catching.</p>
<p><a href="https://metr.org/blog/2025-06-05-recent-reward-hacking/">METR’s research on real-world reward hacking (June 2025)</a> found something more unsettling: training models to avoid <em>detectable</em> cheating sometimes caused them to find <em>less detectable</em> cheating strategies. The field doesn’t have a definitive solution.</p>
<p>The LLM-judge pattern mitigates this. You’re not asking the judge to pass or fail — you’re asking it to score improvement. That’s a much harder goal to game. And cross-model judging means the judge has no incentive to rationalise the output it’s evaluating.</p>
<p>It’s a mitigation, not a guarantee.</p>
<h2>Risk 3: Skills go stale and nobody notices</h2>
<p>Code changes. Standards evolve. A skill written for a codebase with one set of conventions may become actively misleading as the codebase matures. Unlike code, skills don’t have type checkers to tell you when they’re wrong. They’ll silently produce incorrect output.</p>
<p>If you’ve read the <a href="https://blog.dicko.dev/posts/the-documentation-graveyard/">Documentation Graveyard</a> post, this is the same failure mode — knowledge encoded in a format that has no maintenance forcing function. The difference is that skills can be tested. If you maintain your code-before examples alongside the skill, a stale skill will eventually fail its tests.</p>
<p>The discipline required: update the code-before examples when the codebase changes. That’s it. That’s the whole thing. And it’s the thing most teams won’t do.</p>
<h2>Risk 4: The selection mechanism is a black box</h2>
<p>There is currently no telemetry available from Cursor or most other IDEs that tells you whether a specific skill was selected for a given agent session. You can observe the output, but not the mechanism. If a skill isn’t being selected, you won’t know it — you’ll just see no improvement and have to diagnose why.</p>
<p>Your CI pipeline can verify skill selection in controlled runs. In real engineering sessions, you’re partially flying blind. Test as many coding agents as you can in your pipeline, we do a parallel matrix in GitLab CI and test a bunch.</p>
<h2>Risk 5: Agents constrained badly become unhelpfully narrow</h2>
<p>A story from a FOSS Asia presentation, worth repeating: a team built a system where autonomous agents are triggered automatically on master branch CI failures to investigate and fix flaky tests. The system worked — but without an explicit constraint, the agent would identify a failing test, decide to refactor the surrounding code, delay the critical merge, and introduce unrelated risk. The team added a single constraint to their skill: “no unnecessary refactoring.” The agent stayed scoped to the job it was given.</p>
<p>Constraints in skills work both ways. The goal isn’t just to encode what you know — it’s to encode what the agent <em>shouldn’t</em> do in the process of applying what it knows.</p>
<h2>Measuring Whether It’s Working</h2>
<p>If you can’t measure the improvement, you’re running on faith.</p>
<p>The metric we use is first pass success rate: specifically, the proportion of agent sessions that require only a single pass versus multiple passes (re-prompting). Each agent invocation is a pass. A gap of more than five minutes between passes signals the engineer moved on to a new task — that boundary defines a session. One pass in a session is first-time-right. Two or more is iteration.</p>
<p>This matters because each extra pass is directly measurable time and compute cost. Token spend is logged per invocation with timestamps, requiring no new instrumentation. If skills reduce passes, the dollar impact is calculable immediately.</p>
<p>The known blind spot: engineers who fix the agent’s output manually rather than re-prompting don’t appear as multi-pass sessions. Based on our internal survey data, roughly 25% of engineers prefer to fix manually. That means the true re-work rate is likely higher than the session data shows. We acknowledge this explicitly — the survey provides a partial offset, but the metric understates the problem.</p>
<p>The experiment compares a treatment group (engineers with the skill library active) against a control group (same roles, no skills). A meaningful reduction in multi-pass sessions, sustained over roughly two sprints, is the signal we’re looking for. The direction matters more than the absolute number. If skills don’t move the needle, the post-experiment diagnosis looks at: whether skills were actually being selected (testable via the pipeline), whether the skills addressed the right categories of re-prompting, and whether the sample window was long enough.</p>
<p>This is the hypothesis-first, measurement-first discipline described in <a href="https://blog.dicko.dev/posts/i-dont-care-what-you-build-and-neither-should-you/">I Don’t Care What You Build</a>. Articulate what you expect the skill to change before you write it. If you can’t, you’re not ready to write it.</p>
<h2>The Doomsday Version</h2>
<p>I want to end with the failure mode that keeps me honest.</p>
<p>We built an internal app — an MR tracker, AI-generated summaries for managers — by iterating with an AI assistant without any skills. The code worked. The app worked. But it was written in a dialect only the LLM could read fluently. The component structure bore no relationship to the rendered UI. The API calls were in the wrong places. The error handling was symptom-patching all the way down. The engineers who looked at it later understood none of it.</p>
<p>That’s manageable for an internal tool.</p>
<p>The version I think about at 2am: imagine that codebase is production. Imagine it’s 2am, there’s an incident, engineers are in the war room on level 7, and the system they’re trying to debug was vibe-coded by an AI without tribal knowledge. They open the file. They can’t read it because the rely on AI to not only write the code but understand it too. They call the AI for help. The AI gets stuck in a loop of, fix X, oh that didn’t work, I’ll fix Y, oh that didn’t work, I’ll fix X. You’ve seen it before right? What do we do now?</p>
<p>Skills won’t prevent every version of this. But they’re the difference between an AI that generates code that <em>your engineers can read, reason about, and take ownership of</em> — and one that generates code that only the AI understands, which is not ownership culture, it’s a different kind of vendor lock-in.</p>
<h2>Where to Start</h2>
<p>If you’re going to do one thing: write one skill. Pick the piece of knowledge your team repeats most often in code review or advice to juniors (not everything is in PR review comments, there’s so much more that&#39;s verbal) — the comment that appears so regularly it’s practically a meme. That’s your first skill. Write it down the way you’d say it to a new engineer on their first day.</p>
<p>Start with skills in your repo to fast iterate, then expand to mcp distribution later with a proper pipeline and test.</p>
<p>Then set up promptfoo and write tests for it. The discipline of writing the test will force you to articulate exactly what you’re asserting, which will make the skill itself more precise.</p>
<p>That’s a morning’s work. It’s also the beginning of a skill library that actually holds up under the pressure of a production codebase.</p>
<p>The rest is iteration.</p>
<p><strong>References:</strong></p>
<ul>
<li><a href="https://github.com/promptfoo/promptfoo">Promptfoo (open-source, MIT-licensed)</a>- <a href="https://arxiv.org/abs/2510.20270">ImpossibleBench — LLM agents and test exploitation (arXiv:2510.20270)</a>- <a href="https://metr.org/blog/2025-06-05-recent-reward-hacking/">METR: Recent Frontier Models Are Reward Hacking (June 2025)</a>- <a href="https://blog.dicko.dev/posts/the-documentation-graveyard/">The Documentation Graveyard — Beer and Servers Don’t Mix</a>- <a href="https://blog.dicko.dev/posts/i-dont-care-what-you-build-and-neither-should-you/">I Don’t Care What You Build — Beer and Servers Don’t Mix</a>- <a href="https://blog.dicko.dev/posts/the-prototype-that-became-production/">The Prototype That Became Production — Beer and Servers Don’t Mix</a>- <a href="https://blog.dicko.dev/posts/the-question-that-changes-everything/">The Question That Changes Everything — Beer and Servers Don’t Mix</a></li>
</ul>
]]></content:encoded>
  </item>
  <item>
    <title>Teaching Engineers to Speak in Constraints</title>
    <link>https://blog.dicko.dev/posts/teaching-engineers-to-speak-in-constraints/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/teaching-engineers-to-speak-in-constraints/</guid>
    <pubDate>Thu, 02 Apr 2026 12:46:11 GMT</pubDate>
    <category>engineering-management</category>
    <category>leadership</category>
    <category>software-engineering</category>
    <category>software-development</category>
    <category>agile</category>
    <description>Or: Why “We Can’t Do That” and “Sure, No Problem” Are Both Lies</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/teaching-engineers-to-speak-in-constraints/cover.webp" alt=""></figure>
<p>It was 2:15 on a Wednesday afternoon, and Bruce had just killed the room. The business lead from the Australian side was mid-sentence — something about a new ticketing integration, a partnership that could double our event throughput — when Bruce leaned back in his chair, arms crossed, and said the two words that make stakeholders’ eyes glaze over: “Can’t be done.” The pedestal fan in the corner of our Phaya Thai office hummed into the silence. On the other end of the video call, you could hear someone in Melbourne quietly exhale. The conversation was over before it started.</p>
<p>Except it wasn’t. Because the thing Bruce was trying to say — the thing he <em>meant</em> — was right. We couldn’t do what they were asking, the way they were asking for it, in the time they wanted it. He had the instinct. He just didn’t have the sentence.</p>
<p>I leaned forward. “It can be done — but here’s what you’d need to give us, and it’ll be a little slower than you want. Only a little slower.” The room shifted. The Melbourne exhale turned into a question. Options appeared where a wall had been. We left that meeting with a plan that actually worked.</p>
<p>Bruce didn’t need more courage that day. He needed better language. And that distinction — between the instinct to push back and the skill of pushing back <em>well</em> — is what separates engineering teams that get steamrolled from engineering teams that shape their own commitments.</p>
<p>I’ve <a href="https://blog.dicko.dev/posts/your-engineers-cant-negotiate-and-its-your-fault/">written before</a> about the constraints-and-leverage framework — the sentence that changes the conversation in the room. “With [current reality], we can deliver [option A]. To deliver [option B], we’d need [specific ask]. Which would you prefer?” That post was about the skill. This one is about what happens after: how you get an entire team speaking that language, not just the one person who stumbled into it through enough scar tissue.</p>
<h2>When the Map Doesn’t Exist Yet</h2>
<p>Before we go further, let’s name something important: you don’t need to turn every standup into a negotiation exercise.</p>
<p>Core product work — the features your team has built variations of before, in a domain they know, with estimation patterns they trust — doesn’t require the same intensity of constraint-talk. A two-week sprint on familiar ground has predictable scope. You estimate, you commit, you deliver. Constraint negotiation is useful there, but it’s not life-or-death.</p>
<p>Where this skill becomes non-negotiable is when the territory is new. Strategic projects. Experimental scope. The kind of work where the estimate is a guess, nobody knows what the thing will cost until they’ve started building it, and the gap between “what we promised” and “what’s actually possible” can swallow a quarter whole.</p>
<p>That’s the context where the <a href="https://blog.dicko.dev/posts/your-engineers-cant-negotiate-and-its-your-fault/">Lamborghini story</a> lives — low budget, uncharted territory, massive upside if it works. That’s where “sure, no problem” is a death sentence disguised as cooperation, and “can’t be done” is a missed opportunity dressed up as honesty.</p>
<p>As the UCLA basketball coach John Wooden put it, “It’s what you learn after you know it all that counts.” Most engineers know their technical domain. What they haven’t learned is how to translate that knowledge into a conversation that gives the business real options instead of binary outcomes.</p>
<p>The distinction matters because it calibrates the reader’s expectations. If you’re leading a team doing routine feature work in a stable domain, you don’t need a teaching programme. You need good estimation practices and a <a href="https://blog.dicko.dev/posts/the-solutioneering-trap-why-your-best-engineers-are-solving-the-wrong-problems/">healthy relationship with your product counterpart</a>. But when you’re staring at a quarter of unknowns — a new market, a platform migration, a technology bet that could reshape your architecture — that’s when the team that can negotiate the gap honestly will outperform the team that says yes and prays.</p>
<h2>The Skill That Disappears as You Grow</h2>
<p>Here’s something nobody talks about: in startups, constraint-talk happens naturally.</p>
<p>When you’re five engineers in a room with no PM layer, you’re talking directly to the business. There’s no intermediary to absorb the translation. You say “we can’t do X for that budget, but here’s what we can do,” because there’s literally nobody else to say it for you. The founder asks for something impossible, you explain why it’s impossible, you offer what <em>is</em> possible, and you move on. It’s not a skill. It’s survival.</p>
<p>The skill atrophies as organisations grow. PM layers appear. Product managers absorb the constraint conversation. Engineers stop being in the room where commitments are shaped. And eventually, nobody remembers that engineers were ever supposed to be in that conversation at all.</p>
<p>As the management theorist Peter Drucker observed, “The most important thing in communication is hearing what isn’t said.” When your engineers stop speaking in constraints, what’s not being said is every technical reality that should be shaping the commitment. The PM filters it. The estimate absorbs it. And the team delivers something that was never honestly scoped in the first place.</p>
<p>This post is about reversing that atrophy — not by removing intermediaries, but by teaching engineers at any level that the platform for constraint negotiation already exists and they’re allowed to use it.</p>
<h2>The Teaching Playbook</h2>
<p>Most engineering organisations have the same experience with constraint literacy: one senior engineer figures it out through trial and error, becomes the person stakeholders love working with, and everybody else continues defaulting to hero mode or wall mode. The skill doesn’t propagate. It stays locked in one person’s head.</p>
<p>That’s not a talent gap. It’s a coaching gap. And closing it doesn’t require a workshop, a training deck, or a Confluence page that nobody reads. It requires a specific, repeatable pattern — one that works the same way engineers learn any complex behaviour: by watching it, practising it with support, then doing it independently.</p>
<h2>Step 1: The 1:1 Critique</h2>
<p>It starts after a meeting. Not during — after. The meeting where an engineer stayed silent when they should have spoken up, or said “sure, no problem” when the honest answer was “sure, but not all of it.”</p>
<p>In the 1:1, one question: “Why didn’t you say that?”</p>
<p>Not punitive. Curious. The answer is almost always one of two things: “I didn’t realise I was supposed to” or “I didn’t feel like it was my place.” Both reveal the same gap — nobody ever told them this was part of their job. Nobody ever said: when you see a constraint that will affect the commitment, it’s your responsibility to name it. Not the PM’s. Not the tech lead’s. Yours.</p>
<p>The 1:1 is where you set that expectation explicitly. You’re not teaching them what to say yet. You’re teaching them that saying something is expected.</p>
<h2>Step 2: Modelling in the Meeting</h2>
<p>Next time the situation arises, you don’t do it <em>for</em> the engineer. You do it <em>with</em> them.</p>
<p>The technique is specific. You bring the engineer into the conversation with a framing that makes it easy for them to own the response.</p>
<p>“Look, we can’t agree to that — Somchai, wouldn’t you agree?” Then you pause. That pause is the key. You’re not answering for Somchai. You’re creating space for him to wake up and enter the discussion. The framing gives him the answer shape — all he has to do is confirm and elaborate.</p>
<p>Or: “Namfon, that’s the most we could get done for this, right? Like you were saying — anything more is beyond the team’s current capacity.” Pause again.</p>
<p>The engineer owns the response and the commitment, but the manager has demonstrated how to frame it. The constraint is stated. The language is modelled. And the engineer has just experienced what it feels like to say it out loud in a room with stakeholders — and survived.</p>
<p>This is where the <a href="https://blog.dicko.dev/posts/the-design-handoff-is-a-lie-why-your-engineers-and-designers-are-playing-telepho/">design-engineering dynamic</a> comes alive too. The Namfon story from the design handoff post — “What’s the hardest part about building this?” — is constraint-talk in action. The question surfaced a real structural constraint that transformed the estimate. That question only gets asked when someone has been taught it’s their job to surface what’s hard, not just absorb it.</p>
<h2>Step 3: The Follow-Up</h2>
<p>After the meeting, one more conversation: “Next time, I want you to be the one to say it. You saw how it landed. You don’t need me to set it up.”</p>
<p>This is the deliberate handoff from modelling to ownership. The engineer has seen the sentence. They’ve spoken a version of it with scaffolding. Now the expectation is clear: next time, it’s theirs.</p>
<h2>Step 4: Watch for the Unprompted Moment</h2>
<p>The shift has happened when the engineer states a constraint in a meeting without being prompted. No setup. No leading question. Just: “With our current capacity, we can deliver X but not Y — what would you prefer?”</p>
<p>The first time Somchai says that without anyone teeing it up — that’s the signal. Recognise it. Name it. Make it visible to the team. This is <a href="https://blog.dicko.dev/posts/dont-just-eat-and-complain-why-engineering-culture-is-your-best-investment/">recognition as a safety signal</a> — what gets praised publicly defines what gets repeated.</p>
<h2>Why This Works</h2>
<p>It’s not training. It’s apprenticeship. The same way an engineer learns to debug a distributed system or review an architecture proposal — by watching someone do it, doing it with support, then doing it alone.</p>
<p>The 1:1 sets the expectation. The meeting provides the model. The follow-up creates accountability. And the recognition of the unprompted moment makes it stick.</p>
<h2>What Kills Constraint-Talk</h2>
<p>You can teach this skill perfectly and still watch it die. Organisational behaviour is more powerful than individual coaching, and there are five patterns that destroy constraint literacy faster than you can build it.</p>
<p><strong>The disappointed sigh.</strong> The VP who exhales audibly when told “we can’t do all of this” teaches every engineer in the room that honest constraint-setting has a social cost. It doesn’t need to be a reprimand. Body language is enough. One sigh undoes a month of coaching.</p>
<p><strong>“Can’t you just…?”</strong> This phrase reframes a structural reality as an individual failing. It implies the constraints aren’t real — that the engineer is just not trying hard enough. The correct response — “I can explain what that would require, and we can decide together if it’s the right trade-off” — takes courage that most engineers haven’t been trained to have.</p>
<p><strong>Rewarding heroes.</strong> If the engineer who worked weekends to deliver the impossible gets a public shout-out while the engineer who set realistic constraints and delivered on time goes unrecognised, you’ve just trained every person in the room to be a hero. <a href="https://blog.dicko.dev/posts/your-engineers-cant-negotiate-and-its-your-fault/">Hero culture</a> is the enemy of sustainable engineering. It optimises for the dramatic at the expense of the reliable.</p>
<p><strong>Ambiguous commitments.</strong> When nobody explicitly commits to a scope and the meeting ends with “let’s do our best” — that’s hero-mode breeding ground. Ambiguity creates the implicit expectation of “everything,” and engineers absorb the impossible because nobody told them not to. As the economist Thomas Sowell wrote, “There are no solutions, only trade-offs.” Every commitment is a trade-off. If the trade-off isn’t named, it still exists — it just gets made by whoever runs out of hours first.</p>
<p><strong>Constraint-washing.</strong> When the organisation <em>says</em> it values honest estimation and constraint-setting but <em>acts</em> as if it expects everything delivered on time regardless — the say/do gap destroys trust. Engineers learn to perform constraint-setting for the optics while absorbing the impossible for survival. This is worse than no constraint-setting at all, because it adds cynicism to exhaustion. It’s the <a href="https://blog.dicko.dev/posts/the-prototype-that-became-production/">prototype that became production</a> of organisational culture — something temporary and performative that somehow became permanent.</p>
<h2>You Don’t Need Authority — You Need a Platform</h2>
<p>If you’re reading this thinking “great advice for a Director, but I’m a mid-level engineer” — this section is for you.</p>
<p>You don’t need to be a VP to have this conversation with product. You need a platform. And the platform already exists.</p>
<p>Sprint planning. Backlog refinement. Quarterly planning. Design reviews. These are all moments where an engineer at <em>any</em> level can say:</p>
<p>“You want X. With the resources we have now, I can give you Y. If you want X in full, it’ll cost an extra engineer in the team for a quarter, can we trade of some value in another team for this? and move an engineer”</p>
<p>That sentence works whether you’re a tech lead or a first-year engineer. The content changes — a <a href="https://blog.dicko.dev/posts/the-senior-engineer-plateau/">senior engineer</a> quantifies the cost more precisely, frames the trade-off more strategically — but the structure is available to anyone. It also puts the ball back in the court of the person that is responsible for “business value”, the PO, for looking at head count moves. you could go one step further and do this yourself, but that’s the difference betwen good and amazing.</p>
<p>The <a href="https://blog.dicko.dev/posts/why-your-engineers-are-still-waiting/">ownership gradient</a> maps directly here. Station 1 engineers — the ones still waiting for instructions — don’t speak in constraints because they don’t know it’s expected. Station 2 engineers flag constraints but don’t frame them as options. Station 3 and above do constraint-talk naturally, because they’ve learned that shaping commitments is part of their job, not a privilege granted by title.</p>
<p>The coaching mechanism above is how you move people up that gradient. The 1:1 makes the expectation explicit. The modelling shows what it looks like. The follow-up transfers ownership. And the recognition makes it stick.</p>
<h2>Beyond the Planning Meeting</h2>
<p>Constraint-talk doesn’t stop at sprint planning. Once engineers have the language, it applies everywhere the scope can shift.</p>
<p><strong>Scope creep mid-sprint.</strong> “We can add this, but something else needs to come out. What would you like to trade?” This isn’t pushback. It’s a question that assumes the stakeholder is a rational adult who can make trade-off decisions if given real options.</p>
<p><strong>Cross-team dependencies.</strong> “Our team can support your request, but our current sprint has us committed through the 15th. We can start on the 16th, or we can discuss pulling someone off [specific item] to start sooner. Which works better for your timeline?” This is the <a href="https://blog.dicko.dev/posts/breaking-engineering-silos-the-uncomfortable-truth-about-collaboration/">anti-silo language</a> — negotiating across team boundaries requires shared objectives and honest constraints. Without both, cross-team requests either get stonewalled or silently absorbed.</p>
<p><strong>Architecture decisions.</strong> “What’s the simplest way we could do this?” is itself a constraint question — it forces the team to name what can be <a href="https://blog.dicko.dev/posts/the-question-that-changes-everything/">removed rather than added</a>. Every feature request is implicitly a constraint negotiation between what’s desired, what’s feasible, and what’s sustainable.</p>
<p><strong>The estimation conversation.</strong> This is where constraint-talk connects to <a href="https://blog.dicko.dev/posts/the-question-that-changes-everything/">the question that changes everything</a>. “How can we do this in a more simple way?” isn’t a simplification request — it’s a constraint-reframing exercise. It asks the engineer to identify what’s essential versus what’s assumed, and to surface the trade-offs that make the thing actually buildable.</p>
<h2>The Skill That Scales</h2>
<p>Let’s go back to Bruce. He had the instinct to push back. He’d been around long enough to know when something wasn’t going to work. But he opened with a wall — “can’t be done” — because nobody had ever shown him what the alternative looked like. Nobody had modelled the “it can be done, but” sentence in front of him and then handed it to him to try.</p>
<p>Once he had it — once he saw what options opened up in that room when the wall became a door with conditions — he used it for the rest of his career. Not because he attended a workshop or read a book about negotiation. Because he experienced, in a specific moment, the difference between shutting a conversation down and reshaping it. And someone was there to name what had just happened.</p>
<p>The question for engineering leaders is this: how many people on your team right now have the right instinct but the wrong sentence? People who can see the constraint, who know the commitment is unrealistic, who feel the gap between what’s being asked and what’s possible — but stay quiet because nobody ever told them that speaking up was part of the job?</p>
<p>The coaching pattern works. Set the expectation in the 1:1. Model the language in the meeting. Create the pause that lets someone own it. Follow up and hand it off. Watch for the unprompted moment — and when it comes, recognise it loudly, because that’s the moment the skill stops being yours and starts being theirs.</p>
<p>As Daniel Kahneman wrote in <em>Thinking, Fast and Slow</em>, we systematically underestimate the time, cost, and risk of everything we plan. That’s not a bug in human cognition — it’s the default setting. Constraint-talk is how you override the default. Not once, in a planning meeting. Continuously, as a team capability, embedded in how your engineers think and speak.</p>
<p>The <a href="https://blog.dicko.dev/posts/the-org-chart-trap-why-your-company-structure-is-silently-killing-velocity/">org chart</a> won’t teach them this. The process won’t teach them this. The Confluence page definitely won’t teach them this. You will — one meeting, one pause, one follow-up at a time.</p>
<p>Now, if you’ll excuse me, I have a 1:1 in ten minutes with an engineer who said “sure, no problem” in yesterday’s planning session. We need to talk about what “no problem” actually meant — and what they should have said instead.</p>
]]></content:encoded>
  </item>
  <item>
    <title>The Biggest Problem With Our Codebase Is the Developers</title>
    <link>https://blog.dicko.dev/posts/the-biggest-problem-with-our-codebase-is-the-developers/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/the-biggest-problem-with-our-codebase-is-the-developers/</guid>
    <pubDate>Tue, 31 Mar 2026 14:29:51 GMT</pubDate>
    <category>platform-engineering</category>
    <category>developer-experience</category>
    <category>microservices</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <description>Or: What Platform Teams Actually Exist For (And Why Most Get It Completely Backwards)</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/the-biggest-problem-with-our-codebase-is-the-developers/cover.webp" alt=""></figure>
<p>It was a Tuesday afternoon on level 6 when someone asked the question that ended the meeting. Not loudly — just dropped it into a lull in conversation, the way you drop something heavy onto vinyl flooring. The question was: <em>“How many different ways do we handle authentication token expiry across our services?”</em> The whiteboard still had yesterday’s architecture diagram on it. The Thai iced teas had gone warm. Someone started counting on their fingers, then ran out of fingers. The answer, after a few more minutes of uncomfortable archaeology through Slack and Confluence, was eleven. Three of them were wrong. We wouldn’t find out which three for another six months.</p>
<p>That wasn’t a hiring problem. That wasn’t a code review problem. That was a scale problem dressed up as a technical one — and the solution wasn’t going to come from a retrospective or a stronger linting rule.</p>
<blockquote>
<p>“Hell is other people.”* — Jean-Paul Sartre*</p>
</blockquote>
<p>Sartre was writing about existential conflict. He was also, without knowing it, describing what happens when 200 engineers solve the same problem independently.</p>
<p>Here’s the uncomfortable truth that platform teams exist to address: developers — given complete freedom — will make individually rational choices that collectively produce chaos. Not because they’re careless. Not because they’re bad engineers. Because local optimisation doesn’t equal global optimisation. Every engineer who solves the A/B testing problem their own way is making a sensible decision. The result across a thousand engineers is five different implementations, three of which are buggy, none of which interoperate, and all of which need to be maintained indefinitely by people who weren’t there when they were written.</p>
<p>The central tension of engineering at scale is this: <strong>freedom is the source of developer joy, and the source of organisational entropy.</strong> Platform teams exist to manage that tension — not by eliminating freedom, but by making the right path the easy path.</p>
<p>The guardrail isn’t a cage. It’s what lets you drive fast.</p>
<h2>What a Platform Actually Is (And Isn’t)</h2>
<p>Most people hear “platform team” and think: Kubernetes. CI/CD pipelines. Cloud infrastructure. That’s one kind of platform — the runtime kind. But the more interesting and underappreciated kind is the <strong>application platform</strong>: the shared libraries, middleware, conventions, and tooling that shape how engineers write code, not just where that code runs.</p>
<p>Think about what happens in a large engineering organisation without one. Four separate teams need client-side load balancing. Hardware load balancers don’t work at the scale you’re operating — too expensive, and the heartbeat mechanism means a dead node causes a segment of traffic to fail for <em>minutes</em> before the decision to remove it propagates. So each team builds their own implementation. Not because they’re careless. Because the organisation hasn’t yet built the thing that should exist to solve this once. Each team locally optimised. The collective result was four implementations, fragmented across the codebase, some of them sharing organically across team boundaries, others silently diverging. Eventually a platform team was formed for this space and the consolidation began — slowly, as migrations always do.</p>
<p>This is what scale looks like without a platform. Not dramatic failure. Just quiet, compounding fragmentation.</p>
<blockquote>
<p>“The strength of the team is each individual member. The strength of each member is the team.”* — Phil Jackson*</p>
</blockquote>
<p>Jackson was talking about the Chicago Bulls. He was also, unknowingly, describing the economics of shared platform infrastructure. The individual engineer is more capable when the team around them has already solved the foundational problems. The platform is the mechanism by which that happens at scale.</p>
<p>The test for whether something belongs in the platform is simple: <strong>would at least five teams build this independently if it didn’t exist?</strong> If yes, the platform should own it. If no, you might be building premature abstraction.</p>
<h2>The Scale Threshold</h2>
<p>Platform teams make no sense below a certain size. Two engineers don’t need a platform — they need a conversation. At ten engineers, shared documentation is probably sufficient. The investment starts paying off around eight to ten teams, and the ROI compounds from there.</p>
<p>Below that threshold, organic sharing works. Engineers naturally cross-pollinate solutions, grab code from each other’s repositories, have corridor conversations. Joel Spolsky once wrote about “the Joel Test” as a quick measure of engineering hygiene — platform investment is something like that, but for organisations rather than teams. You can operate without it for a while. The symptoms are subtle at first.</p>
<p>Above the threshold, organic collaboration breaks down faster than documentation can keep up. The divergence tax compounds. You stop being able to answer basic questions — “how many ways do we do X?” — without an audit. New engineers joining take months to understand “how we do things here” because “how we do things here” is now twelve different answers depending on which team you landed on.</p>
<p>This connects directly to <a href="https://blog.dicko.dev/posts/the-impact-of-paved-paths-and-embracing-the-future-of-development/">the paved paths work we’ve covered in this series</a> — the platform is the delivery mechanism for paved paths at scale. The path is the convention. The platform is the infrastructure that makes following the path easier than not following it.</p>
<h2>The Modular Lesson (That Microsoft Already Learned)</h2>
<p>The original .NET Framework gave you everything. The problem was you took all of it whether you needed it or not — the kitchen sink came standard. One of the central design decisions in .NET Core (later .NET 5+) was a shift to modular, independent libraries with inversion of control at the seams. Don’t want the full HTTP stack? Don’t include it. Need to swap the logging abstraction? Write a new implementation for the interface everything is using already and override it. The framework became composable.</p>
<p>This is the design philosophy platform teams should steal — not the technology, the <em>philosophy</em>. Build modular. Define clear interfaces. Inversion of control at every seam so consumers can override what they need to override. The platform should be composable, not monolithic.</p>
<p>We learned this lesson at Agoda the expensive way.</p>
<p>When breaking up a large monolith, the business domain separation was done well. But the cross-cutting concerns — auth, observability, HTTP pipeline middleware, attribution logic — were bundled into a single shared library. The intent was good: <em>“everything you need, packaged up nicely.”</em> The result was what Joe Armstrong, creator of Erlang, once described as the Gorilla/Banana problem: you wanted a banana, you got a gorilla holding a banana, and the entire jungle along with it.</p>
<p>One initialisation method: AddPlatform(). Inside it, anything could go wrong. When it did, nobody knew how it was woven together.</p>
<p>Two failure modes followed immediately. The first was a debugging black box — “why can’t I debug this locally?” became a constant escalation. When AddPlatform() failed, engineers had no visibility into what was happening inside it. The second was structural: because the HTTP pipeline lived inside the platform library, teams needing to make changes to cross-cutting concerns — attribution logic being the example — had to go through the platform team. Not because the platform team was obstructive. Because the <em>architecture</em> made them the mandatory path. A business logic change became a platform ticket.</p>
<p>The platform team didn’t become a bottleneck because they were slow or territorial. They became one because the design created the dependency.</p>
<p>Microsoft had already figured this out. We learned it again, from scratch, the usual way.</p>
<h2>Divergence Is a Signal, Not a Problem</h2>
<p>Here’s the argument that most platform teams get exactly backwards.</p>
<p>When an engineer goes off the platform path — builds something the platform doesn’t support, rolls their own solution — the instinct is to treat it as a problem to be standardised away. Get everyone back on the path. Raise it in the next architecture review. File a ticket to add it to the roadmap.</p>
<p>This is exactly wrong.</p>
<p><strong>Divergence is innovation.</strong> The engineer who went off-path did so because the path didn’t solve their problem. That’s not insubordination — it’s a signal. They found a gap in the platform. And they’re the best possible person to have found it, because they’re closest to the problem it failed to address.</p>
<p>The right model: track divergence and feed it back into the platform roadmap. Not every divergence becomes a platform feature — some are genuinely team-specific. But the <em>pattern</em> of where teams diverge tells you exactly where the platform is under-invested. It’s free product research. It’s the thing most platform teams would pay for if they could, sitting right there in their own codebase, being treated as a compliance problem.</p>
<p>Practically, this means building the platform with explicit extension points — seams where teams <em>expect</em> to plug in their own behaviour. It means creating a lightweight channel for teams to contribute divergent solutions back into the platform, rather than treating every off-path decision as a governance failure. Martin Fowler describes this as <a href="https://martinfowler.com/bliki/CodeOwnership.html">weak code ownership applied at scale</a> — modules have owners, but the contribution model isn’t closed.</p>
<p>The framework: <strong>platforms that prevent divergence become bottlenecks. Platforms that channel divergence become better platforms.</strong></p>
<h2>The Guardrail vs. The Gate</h2>
<p>A <strong>guardrail</strong> prevents bad outcomes without preventing movement. You can still drive fast. You can still choose your route. The guardrail activates when you’re about to go off a cliff.</p>
<p>A <strong>gate</strong> stops you. Full stop. Someone decides whether you may proceed.</p>
<p>Most platform teams start as guardrails and gradually become gates. The progression is always the same: build shared library → developers use it → edge case emerges → platform team adds a constraint to prevent misuse → another edge case → another constraint → the library now has seventeen configuration options, a ticket process for exceptions, and an approval workflow for new adopters.</p>
<p>You’ve built a gate and called it a platform.</p>
<p>The guardrail version has sensible defaults that cover 90% of cases. For the other 10%, you can override. You don’t need to ask permission. The guardrail is there for the cases where developers are about to make a decision they’ll regret — not for every decision.</p>
<p>There’s a simple test for which one you’ve built: <strong>if using the platform increases a developer’s cognitive load compared to rolling their own solution, the platform has failed.</strong> The whole point is to reduce cognitive load — to make the right thing the easy thing, the default thing. A developer who avoids the platform because it’s more work than the alternative is a developer the platform has lost. And a developer the platform has lost will write their own implementation, and you’re back to eleven ways of handling token expiry.</p>
<p><a href="https://teamtopologies.com/">Team Topologies</a> puts it clearly: the measure of a platform team’s success is how little stream-aligned teams need to think about it. Not how many features the platform has. Not how comprehensive the documentation is. How little friction it creates.</p>
<h2>Platform as Product (The Mindset Shift)</h2>
<p>The most underappreciated insight in platform engineering: <strong>your developers are your customers.</strong> Every instinct you have about product development — understand the user’s job to be done, measure adoption, gather feedback, reduce friction in the onboarding experience — applies here.</p>
<p>Most platform teams operate as internal utilities. They build what they think the organisation needs. They respond to requests. They define roadmaps based on their own assessment of the architecture. This works until it doesn’t — until adoption stalls because the platform solves the wrong problems, until teams work around it rather than with it, until the platform team is busy maintaining things nobody uses.</p>
<p>One quarter, our platform team took on a DX initiative: move local development environments from shared QA servers to local test containers. The motivation was real — if another team deployed a broken package to a shared QA server your developers depended on, your entire team could be blocked for hours waiting for a fix that wasn’t yours to make. Local isolation solved the dependency problem. But it introduced a new one: startup time. A freshly initialised local container is cold. A shared QA server is warm and pre-loaded. Solve one problem badly and you trade a fragile-but-fast environment for a reliable-but-slow one.</p>
<p>The platform team used local development metrics to measure startup time impact and ensure it stayed within acceptable bounds. Not a subjective argument about whether the new setup “felt” slower — data on whether it actually was, and by how much. That’s what allowed the migration to proceed with confidence rather than stalling in a preference debate.</p>
<p>This connects directly to <a href="https://blog.dicko.dev/posts/the-inner-loop-nobody-measures/">the inner loop work we’ve covered</a> — developer experience metrics aren’t just useful for measuring problems. They’re the mechanism by which platform teams are held accountable for the quality of the tools they build.</p>
<p>The product mindset requires:</p>
<ul>
<li><strong>Adoption metrics.</strong> Not just “is the library available” but “are teams actually using it?” And if not, why not?- <strong>Office hours and pairing.</strong> Platform teams that sit in an ivory tower and throw APIs over the wall produce platforms nobody loves. Platform teams that work alongside consuming teams build things that fit.- <strong>Explicit deprecation processes.</strong> Nothing erodes trust faster than a breaking change with a short notice window.- <strong>Developer satisfaction as a first-class outcome.</strong> <a href="https://itrevolution.com/product/accelerate/">Accelerate</a> shows clearly that deployment frequency and lead time for changes are correlated with engineering culture — the platform is a direct input into both.</li>
</ul>
<h2>Platform Teams Must Own a Production System</h2>
<p>This is the mechanism that makes the product mindset actually work, and it deserves naming as a structural policy rather than a vague principle.</p>
<p>Platform teams that don’t own a system that runs on their own platform gradually lose the ability to feel what they’ve built. The setup docs get longer because nobody on the platform team had to follow them recently. The initialisation ceremony grows because nobody was inconvenienced by the complexity. The error messages stay cryptic because nobody on the platform team debugged them under pressure at 5pm.</p>
<p>We had an example of this done right. A platform engineer, working on the team’s own production system, noticed inconsistent error handling patterns across the codebase — different teams handling errors differently, no standardisation, no obvious right answer in the existing platform. Because he was <em>using</em> the platform himself rather than just building it, he saw the gap firsthand rather than hearing about it via ticket. He built an npm package to standardise error handling. Critically: because he’d had to add it himself, there was no complex setup documentation. No onboarding guide. No “read this before you start.” Just: install the package, add one snippet to your page-level components, done. The solution was shaped by the experience of being the first person to use it.</p>
<p>That’s the feedback loop that bad platform teams don’t have. Dogfooding isn’t just a quality assurance mechanism. It’s an empathy mechanism. The engineer who has to add their own error handling library to their own system at 4pm on a Friday will write a better library than the one who specifies it from a design document.</p>
<blockquote>
<p>“In theory, there is no difference between theory and practice. In practice, there is.”* — *<em><strong>old engineering adage</strong></em></p>
</blockquote>
<p>The fix is structural: <strong>the platform team must own and operate a production system that uses the platform.</strong> Not a demo environment. Not a reference implementation. A real system, serving real traffic, that the platform team is on-call for. This is dogfooding, but with teeth.</p>
<h2>The “That’s the Platform Team’s Job” Anti-Pattern</h2>
<p>Name this one explicitly, because it lives in every large organisation and almost nobody talks about it.</p>
<p>A developer identifies a problem — something genuinely useful that the platform doesn’t yet address. They raise it. The response: <em>“That’s the platform team’s job. File a ticket.”</em> The ticket enters the backlog. The developer’s team ships their product deadline, rolls their own solution, and now there are two implementations of the thing. The platform team eventually builds their version, which the developer’s team has no incentive to migrate to because their version already works.</p>
<p>Cost: duplicated effort, inconsistency, developer frustration, and a platform roadmap that doesn’t reflect what teams actually need.</p>
<p>The fix isn’t to tell developers to wait. It’s to build a model where developers <em>can</em> contribute to the platform with appropriate oversight — what Tim O’Reilly coined — and Martin Fowler champions — as the <a href="https://martinfowler.com/bliki/InnerSource.html">innersource model</a>. Not “commit to the platform repo without review,” but “here is the process by which a team with a real problem can extend the platform and have that extension considered for inclusion.”</p>
<p>The caveat: not everything a developer wants to add belongs in the platform. The platform team’s job includes saying no to contributions that are too team-specific, too untested, or at odds with the architecture direction. But the platform team that says no to everything and makes developers file tickets for everything has confused governance with gatekeeping.</p>
<p>This is a structural sibling to <a href="https://blog.dicko.dev/posts/breaking-engineering-silos-the-uncomfortable-truth-about-collaboration/">the silo problem we’ve written about before</a> — the innersource model is one of the few mechanisms that structurally forces cross-team collaboration without requiring anyone to reorganise.</p>
<h2>The Bottom Line</h2>
<p>Platform teams exist because scale breaks the informal coordination mechanisms that work perfectly well at twenty engineers and catastrophically badly at two hundred. They exist not because developers can’t be trusted, but because developers <em>can</em> be trusted to solve every problem in front of them — including ones they shouldn’t each have to solve individually.</p>
<p>The platform’s job is to make the right path the easy path. To absorb the solved problems so teams can focus on the unsolved ones. To reduce the cognitive load of building a new service from “learn twelve different conventions” to “follow this template, extend where you need to.”</p>
<p>The platform teams that fail treat this as an infrastructure problem — something to govern, constrain, standardise. The platform teams that succeed treat it as a product problem — something to design, measure, iterate, and ultimately make their customers love.</p>
<blockquote>
<p>“You can design and create, and build the most wonderful place in the world. But it takes people to make the dream a reality.”* — Walt Disney*</p>
</blockquote>
<p>Your engineers are the people who make the architecture real. The platform exists to give them the best possible conditions to do that. If you’re building a gate and calling it a guardrail, you’re not protecting the codebase — you’re just adding a tollbooth to the road.</p>
<p>As for me — I’m going to go check how many different ways we currently handle retry logic across our services. I have a feeling the answer is going to require more than one hand.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Introducing DX Telemetry Manager</title>
    <link>https://blog.dicko.dev/posts/introducing-dx-telemetry-manager/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/introducing-dx-telemetry-manager/</guid>
    <pubDate>Tue, 24 Mar 2026 01:43:17 GMT</pubDate>
    <category>developer-experience</category>
    <category>open-source</category>
    <category>devops</category>
    <category>software-engineering</category>
    <category>software-development</category>
    <description>Or: An Open Source Dashboard for the Developer Inner Loop</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/introducing-dx-telemetry-manager/cover.webp" alt=""></figure>
<p>There’s a particular satisfaction to the moment a dashboard shows you something you didn’t know you were missing.</p>
<p>It was somewhere around the third week after the first rollout. We were looking at the clientside build data — not for any specific reason, just following the numbers — when three engineers showed up as consistent outliers. Nearly double the median webpack build time. Every day. Reproducibly. We pulled the hardware correlation. Seventh generation Intel chips. Laptops that hadn’t been renewed in three years. Nobody had filed a ticket. Nobody had flagged it in a retro. The engineers themselves had adapted to it so completely that slow had become their normal. We immediately got them replacement laptops.</p>
<p>We had built a tool to measure local build times. It found a hardware equity problem. That’s what happens when you make the invisible visible: you find the thing you didn’t know to ask about.</p>
<p>We’ve now open-sourced everything — the dashboard, the collection clients, all of it. This post is about what it is, how it works, and how to get your own data flowing.</p>
<p>If you want the full story of <em>why</em> the inner loop matters and what we found when we started measuring it properly, start with <a href="https://blog.dicko.dev/posts/the-inner-loop-nobody-measures/"><em>The Inner Loop Nobody Measures</em></a> and <a href="https://blog.dicko.dev/posts/what-we-actually-instrumented-and-what-we-found/"><em>What We Actually Instrumented — and What We Found</em></a>. If you’re here for the tool, read on.</p>
<h2>What This Is</h2>
<p><a href="https://github.com/agoda-com/Local-Dev-Telemetry-Manager">Agoda DevExTelemetry</a> is an open-source dashboard that collects, stores, and visualises developer inner loop telemetry — build times, startup times, HMR times, test execution rates — across .NET, JavaScript/TypeScript, and JVM stacks. It’s not what we use, we use our data platform on prem, which is slightly hard to open source :) so I built this so small companies can run something simple on their infrastrucutre to get going faster with monitoring dev local metrics.</p>
<p>It is specifically <em>not</em> a CI monitoring tool. CI tells you about the outer loop: clean builds, controlled environments, hardware nobody develops on. This tells you what the engineer at the laptop actually experiences, every debug cycle, every day. Those are different measurements. As we learned with Somchai’s five-minute debug cycle hiding behind a 30-second compile metric, the gap between them is where the real problems live.</p>
<p>The system has three parts:</p>
<ul>
<li><strong>Collection clients</strong> — build and test reporter plugins that run automatically alongside normal development, zero friction for engineers- <strong>A central ingest API</strong> — receives timing and hardware data from client machines and stores in sql-lite- <strong>A dashboard</strong> — makes the data visible, queryable, and comparable across teams and over time from teh same database
All three are open source. The clinet collect and part of teh ingestion API are running across roughly 700 engineers at Agoda. The ingestion API is “based” on the API we use internally at Agoda, and we’ve done some of the same performance optimiztions on it as well.</li>
</ul>
<h2>Architecture</h2>
<p>Simple by design. The goal was something any team could deploy without dedicated infrastructure or a platform team standing it up.</p>
<ul>
<li><strong>Backend:</strong> .NET 10, ASP.NET Core (Kestrel), EF Core with SQLite- <strong>Frontend:</strong> React 19, TypeScript, Vite 8, Tailwind CSS 3, Recharts- <strong>Deployment:</strong> GitHub Actions → Azure App Service (Linux); or self-host anywhere .NET runs. You can take the artifact or docker image and use it yourself inside your company.
SQLite keeps the operational footprint minimal. No separate database to provision, no connection strings to manage, no infrastructure dependency for most team-scale deployments. It’s one of those architectural choices that looks aggressively simple until you realise that simple is the point — the goal is to remove every possible reason not to deploy it.</li>
</ul>
<h2>What Gets Collected</h2>
<p>The ingest API accepts telemetry from all three stacks via dedicated endpoints:</p>
<p>Endpoint What it receives POST /dotnet .NET compile time, ASP.NET startup time, time to first response POST /dotnet/nunit NUnit / xUnit test results and durations POST /webpack Webpack full build time and HMR time POST /vite Vite full build time and HMR time POST /jest Jest test results POST /vitest Vitest test results POST /junit JUnit test results POST /scala/scalatest ScalaTest results POST /gradletalaiot Gradle Talaiot build metrics</p>
<p>The .NET endpoint collects three separate metrics — compile, startup, and first response — because each degrades independently. Combining them into a single &quot;build time&quot; would have hidden the Somchai problem entirely. Startup finishing doesn&#39;t mean the service is ready. First response time is the number that actually matters to the engineer staring at the browser.</p>
<p>The hardware metadata collected alongside build timings is what made the laptop correlation possible. Without knowing what machine each build ran on, fast and slow are just numbers with no explanation.</p>
<h2>The Collection Clients</h2>
<h2>.NET</h2>
<pre><code>dotnet add package Agoda.Builds.Metrics
dotnet add package Agoda.DevFeedback.AspNetStartup
dotnet add package Agoda.Tests.Metrics.NUnit   # or .xUnit
</code></pre>
<p>The MSBuild plugin hooks into the build pipeline automatically — no registration, no config file, no ceremony. The ASP.NET package instruments startup time and time to first response as separate metrics. After the next build, data starts flowing.</p>
<p>The right home for these packages is your <a href="https://blog.dicko.dev/posts/starting-a-paved-path-with-net-templates/">paved path .NET template</a>. Wire them in at the template level and every new service comes instrumented by default. Engineers don’t need to think about it.</p>
<h2>JavaScript / TypeScript</h2>
<pre><code>npm install agoda-devfeedback-vite2
# or
npm install agoda-devfeedback-webpack
</code></pre>
<p>Add the plugin to your Vite or webpack config — one line. HMR timing starts flowing alongside full build time. HMR is the metric CI has never seen and never will: it only exists on local machines, it degrades silently as dependency graphs grow, and it has an outsized effect on frontend developer experience. When HMR drifts from 200ms to 800ms over three months of normal feature development, no pipeline metric catches it. This does.</p>
<p>If you’re running the <a href="https://blog.dicko.dev/posts/bridging-worlds-making-net-bff-and-reactvite-play-nice-in-development/">.NET BFF with Vite setup</a>, the Vite plugin drops in cleanly alongside the existing config.</p>
<h2>JVM</h2>
<pre><code># Gradle Talaiot plugin, JUnit, ScalaTest clients
# github.com/agoda-com/java-local-metrics
</code></pre>
<p>JVM coverage exists and is open source, but it’s the least mature of the three stacks — it hasn’t received the same investment as the .NET and JS clients. If you work primarily in Kotlin, Scala, or Java, contributions are open and genuinely welcome. The JVM ecosystem deserves the same quality of inner loop telemetry as .NET and JavaScript.</p>
<h2>Pointing the Clients at Your Deployment</h2>
<p>Two options, depending on how much configuration you want to manage per machine.</p>
<p><strong>Option A — environment variable:</strong></p>
<pre><code>export DEVFEEDBACK_URL=https://your-devex-telemetry.example.com
</code></pre>
<p>Set it in your shell profile or your team’s standard dev environment setup script. Data flows on the next build.</p>
<p><strong>Option B — internal DNS:</strong></p>
<p>Create a DNS record that resolves to your deployment. Zero per-machine configuration. Preferred for larger teams where you want the telemetry to be truly invisible — engineers don’t set anything up, the data just arrives.</p>
<p>The clients default to <a href="http://compilation-metrics">http://compilation-metrics</a> make this resovle with your local DNS to where ever you host the API.</p>
<p>Either way, from installation to first data point is measured in minutes, not days.</p>
<h2>What the Dashboard Shows</h2>
<p>Three views, corresponding to the three inner loop problem areas.</p>
<p><strong>API Build Performance</strong> — compile time, startup time, and time to first response for .NET services, charted separately. This is where the Somchai story lives visually: a compile metric that looks fine sitting next to a first response time that tells a completely different story. P50, P75, P90 breakdowns over time, filterable by service and by team.</p>
<figure><img src="https://blog.dicko.dev/posts/introducing-dx-telemetry-manager/image_2.webp" alt="" loading="lazy"></figure>
<p><strong>Clientside Build Performance</strong> — HMR time versus full build time for webpack and Vite, with hardware correlation available. This is the view that found the laptop problem. When three engineers show up as consistent outliers and you can correlate their build times against their hardware specs, the chart does the work that no retro conversation ever would.</p>
<figure><img src="https://blog.dicko.dev/posts/introducing-dx-telemetry-manager/image_3.webp" alt="" loading="lazy"></figure>
<p><strong>Test Run Performance</strong> — pass rates, durations, per-suite and per-test drill-down, and — critically — local versus CI execution rate comparison. The execution rate view is where you find the test suites that are green in CI and abandoned in practice. If a suite has a 15% local execution rate on days with active development, that’s not an engineering discipline problem. That’s a friction problem. The data tells you which suites, and the comparison against CI tells you the gap.</p>
<figure><img src="https://blog.dicko.dev/posts/introducing-dx-telemetry-manager/image_4.webp" alt="" loading="lazy"></figure>
<h2>Ownership, Not Surveillance</h2>
<p>One thing worth naming directly, because it comes up: the dashboard doesn’t tell teams what to do. It makes data visible. What teams do with it is their call.</p>
<p>This follows the same model as production monitoring. You wouldn’t centralise incident response into a platform team and have product teams ignore their own Grafana dashboards. The inner loop deserves the same ownership structure: the team that built the service is the team that runs it, and the team that runs it should own what their developer experience actually looks like. The platform provides the tooling. Teams own the signal.</p>
<p>The teams that have engaged with this most have done so because the data gave them something concrete to take to a conversation — not “our builds feel slow” but “our first response time is 4 minutes 47 seconds at P75 and three months ago it was 40 seconds.” That’s a number with a history. It creates accountability without requiring anyone to feel accused.</p>
<p>The <a href="https://blog.dicko.dev/posts/semantic-monitoring-the-question-youre-not-asking-about-your-production-systems/">Semantic Monitoring post</a> covers the same philosophy applied to production systems. This is the development-side complement: the same argument, one loop earlier in the cycle.</p>
<h2>Get Started</h2>
<p><strong>Repos:</strong></p>
<ul>
<li><strong>Dashboard:</strong> <a href="https://github.com/agoda-com/Local-Dev-Telemetry-Manager">https://github.com/agoda-com/Local-Dev-Telemetry-Manager</a>- <strong>.NET clients:</strong> <a href="https://github.com/agoda-com/dotnet-build-metrics">https://github.com/agoda-com/dotnet-build-metrics</a>- <strong>JS clients:</strong> <a href="https://github.com/agoda-com/devfeedback-js">https://github.com/agoda-com/devfeedback-js</a>- <strong>JVM clients:</strong> <a href="https://github.com/agoda-com/java-local-metrics">https://github.com/agoda-com/java-local-metrics</a>
<strong>Install the clients, point them at a deployment, and data starts flowing on the next build.</strong> The hardest part is genuinely the deploy, not the instrumentation. And for most teams, a single Azure App Service instance on the free or basic tier handles team-scale data without breaking a sweat.</li>
</ul>
<p>If you find something interesting in your data — a correlation we haven’t thought to look for, a problem the tooling surfaces that we didn’t anticipate — open an issue or start a discussion. The laptop finding wasn’t in our design spec. The abandoned test suite finding wasn’t either. Good observability keeps surprising you.</p>
<h2>The Full Series</h2>
<ul>
<li><a href="https://blog.dicko.dev/posts/the-inner-loop-nobody-measures/"><em>The Inner Loop Nobody Measures</em> </a>— why the inner loop is a black box in most engineering organisations, and why that’s a problem worth solving- <a href="https://blog.dicko.dev/posts/what-we-actually-instrumented-and-what-we-found/"><em>What We Actually Instrumented — and What We Found</em> </a>— what we measured, how the collection works, and what the data found that we weren’t looking for- <em>Introducing Agoda.DevExTelemetry</em> — you’re here</li>
</ul>
<h2>Related Reading</h2>
<ul>
<li><a href="https://blog.dicko.dev/posts/bridging-worlds-making-net-bff-and-reactvite-play-nice-in-development/"><em>Bridging Worlds: Making .NET BFF and React/Vite Play Nice in Development</em></a> — the BFF and Vite setup the frontend telemetry instruments- <a href="https://blog.dicko.dev/posts/stop-copying-prod-into-dev-test-data-strategies-that-actually-scale/"><em>Stop Copying Prod Into Dev: Test Data Strategies That Actually Scale</em></a> — what the telemetry found; what this post fixed- <a href="https://blog.dicko.dev/posts/starting-a-paved-path-with-net-templates/"><em>Starting a Paved Path with .NET Templates</em></a> — where the telemetry clients belong in the paved path from day one- <a href="https://blog.dicko.dev/posts/semantic-monitoring-the-question-youre-not-asking-about-your-production-systems/"><em>Semantic Monitoring: The Question You’re Not Asking About Your Production Systems</em></a> — the production-side complement to this series- <a href="https://blog.dicko.dev/posts/mesh-programming-where-visual-design-meets-synchronized-development/"><em>Mesh Programming: Where Visual Design Meets Synchronized Development</em></a> — why inner loop speed directly affects collaborative design-engineering workflows- <a href="https://blog.dicko.dev/posts/the-impact-of-paved-paths-and-embracing-the-future-of-development/"><em>The Impact of Paved Paths and Embracing the Future of Development</em></a> — the broader paved path philosophy this tooling sits inside</li>
</ul>
]]></content:encoded>
  </item>
  <item>
    <title>What We Actually Instrumented — and What We Found</title>
    <link>https://blog.dicko.dev/posts/what-we-actually-instrumented-and-what-we-found/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/what-we-actually-instrumented-and-what-we-found/</guid>
    <pubDate>Mon, 23 Mar 2026 13:15:06 GMT</pubDate>
    <category>developer-experience</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>devops</category>
    <description>Or: Three Engineers, One Hardware Correlation, and the Test Suite Nobody Was Running</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/what-we-actually-instrumented-and-what-we-found/cover.webp" alt=""></figure>
<p>The laptop was three years old. That’s not a detail anyone had written down anywhere.</p>
<p>It sat on the desk the way old laptops do — slightly warmer than it should be, fan spinning at a frequency that had become background noise months ago. The engineer using it had filed no complaints, raised no ticket, mentioned nothing in standup. Builds were slow. Builds were always slow. You learned the rhythm of it: hit run, switch tabs, come back. It was just how things were on that machine, in the same way there was Coffee Mate and not milk in my coffee that morning — it’s just the way things were. You stop questioning the water you swim in.</p>
<p>We found the laptop with data. Not with a conversation, not with a complaint, not with a survey. With telemetry that we’d built to measure something else entirely.</p>
<p>That’s the thing about actually measuring the inner loop: you find problems you didn’t know to look for.</p>
<h2>First, What We Were Trying to Measure</h2>
<p>Post 1 made the case that the inner loop is unmeasured in most engineering organisations — the tight cycle of write, build, test, see that engineers live in all day, every day, while production monitoring stays green and dashboards stay happy. If you haven’t read it, <a href="https://blog.dicko.dev/posts/the-inner-loop-nobody-measures/"><em>The Inner Loop Nobody Measures</em></a> is the place to start. This post is about what we actually did about it.</p>
<p>The goal was straightforward: instrument every meaningful link in the chain. Not just CI build time. Not just compile time. The whole thing, on developer machines, as engineers actually experience it.</p>
<p>The inner loop isn’t one event. For a .NET backend service, it looks like this:</p>
<p><strong>compile → startup → prewarm / first request → browser ready</strong></p>
<p>For a frontend change:</p>
<p><strong>save → HMR → browser update</strong></p>
<p>For a test run:</p>
<p><strong>trigger → Fixture start-up → execution → result</strong></p>
<p>Each link can degrade independently. Each link had, in our case, degraded independently, at different times, for different reasons, completely invisibly. We needed to measure all of them — separately, because aggregating them hides exactly the kind of problem we were trying to find.</p>
<h2>Zero Friction by Design</h2>
<p>Before we get into what we found, a word on how the collection works — because the design choice here matters enormously.</p>
<p>We made a deliberate decision early: the telemetry had to be invisible to engineers. Not opt-in. Not a tool they had to remember to run. Not a script to execute, a config to set up, or a plugin to manually install. Zero friction meant zero friction — the data had to just <em>happen</em>, as a consequence of the normal act of building and testing code.</p>
<p>For .NET, that meant MSBuild and test reporter plugins that hook automatically into the build pipeline, and middleware that&#39;s imported only on debug runs automatically from an IHostingStartup:</p>
<pre><code>dotnet add package Agoda.Builds.Metrics
dotnet add package Agoda.DevFeedback.AspNetStartup
</code></pre>
<p>That’s the entire engineer-facing installation. The MSBuild plugin instruments compile time. The ASP.NET package instruments startup time and time to first response — separately, because as Somchai taught us, those are not the same number. After the next build, data starts flowing.</p>
<p>For frontend, it’s a one-line addition to a Vite or webpack config after installing the package:</p>
<pre><code>npm install agoda-devfeedback-vite2
# or
npm install agoda-devfeedback-webpack
# or
npm install agoda-devfeedback-rsbuild
</code></pre>
<p>Data goes to a central endpoint, configured via a DEVFEEDBACK_URL environment variable — or via an internal DNS record, which means zero per-machine configuration at all for teams where you want it truly invisible. The endpoint is <a href="https://github.com/agoda-com/Local-Dev-Telemetry-Manager"><em>Agoda.DevExTelemetry</em></a>, the open-source dashboard we&#39;ll cover in Post 3. It stores everything and makes it queryable across teams and over time.</p>
<p>Worth being explicit about what we collect: build performance timing and hardware specs. Not what anyone is working on, not file names, not diffs. Just timing and machine metadata. The fact that both the clients and the dashboard are open source means anyone can inspect exactly what gets sent — transparency as a design principle, not an afterthought.</p>
<p>Most engineers don’t know the telemetry is running. Some still don’t. That’s by design.</p>
<h2>What Compile → First Response Actually Looks Like</h2>
<p>The Somchai finding from Post 1 — that a 30-second compile was hiding a five-minute debug cycle — became the clearest example of why splitting the chain matters.</p>
<p>The compile metric was accurate. It was measuring one link. The dashboard shows all three .NET links separately: compile time, ASP.NET startup time, and time to first response. When you put them side by side, the picture is completely different from what any single metric would suggest.</p>
<p>In our case: compile was fast, startup was reasonable, and first response was where everything collapsed. The service was running a web server prewarm routine at startup that pulled real data from QA servers in the data center — cache population, database warm-up, dependency initialisation. All of it legitimate. All of it adding four and a half minutes to every debug cycle. All of it happening in the gap between “startup complete” and “something I can actually click on.”</p>
<p>Once it was visible, it was fixable. The prewarm routine got moved to a background task after the host reported ready. First response time dropped. Engineers stopped needing a ritual to fill the wait time.</p>
<p>The fix took less than a day. Finding it took months of the problem existing invisibly.</p>
<h2>Finding 1: The Laptop Nobody Knew Was Slow</h2>
<p>The hardware correlation finding arrived before we expected it.</p>
<p>Almost immediately after the first rollout of the webpack telemetry, three engineers showed up as outliers in the build data. Not slightly slower — nearly double the median time, consistently, across multiple days and build types. Because the telemetry collected hardware data alongside build performance, we could correlate the two. The three outliers were all running 7th generation Intel chips. Their laptops hadn’t been renewed in three years.</p>
<p>We got them new machines within the week.</p>
<p>The obvious takeaway is that we found and fixed a hardware equity problem. But there’s a less obvious one: those engineers knew their builds felt slow. They’d adapted to it. Without a comparison — without “your webpack time is 1.9x the team median and here is the chart” — there’s no lever to pull. The feeling of “my machine is slow” is easy to dismiss or deprioritise. The data created the justification for action.</p>
<p>There’s also a more uncomfortable version of this finding: how many engineers on your team are running old hardware right now, have adapted to it, and have never said anything because they’ve normalised the experience? You won’t find out with a survey. You’ll find out with data.</p>
<p>As the statistician George Box wrote, “All models are wrong, but some are useful.” Our model of developer productivity had no term for hardware variation because we’d never measured it. Adding the measurement didn’t make the model right — it made it less wrong in a way that mattered.</p>
<h2>Finding 2: The Test Suite Nobody Was Running</h2>
<p>The second finding came from a comparison we almost didn’t think to make.</p>
<p>We were looking at test execution rates — how often given test suites were being run. We had CI data, which told us how often suites ran in the pipeline. We had local data from the test reporter plugins, which told us how often suites ran on developer machines.</p>
<p>For some suites, those numbers were radically different.</p>
<p>Certain test suites were running constantly in CI and almost never on local machines. The local execution rate was close to zero on days with active development on those services. We started asking why.</p>
<p>The answer: Docker Compose with bash scripts and sleeps. Running those tests locally meant dropping out of the IDE, opening a terminal, executing a setup script, waiting for multi-gigabyte containers to spin up, hoping the timing worked, and based on laptop hardware sometimes you would need to vary the sleep time, and then running the tests from your IDE. Engineers had stopped bothering. They’d learned, through experience, that it was faster to just push and let CI handle it. Entirely rational individual behaviour in response to a friction-filled system. Completely invisible unless you’re comparing local execution rates against CI execution rates for the same suite.</p>
<p>Here’s why that matters: when engineers can only run tests in CI, CI becomes part of the inner loop. They write code, push, wait for the pipeline — anywhere from five to forty minutes depending on your setup — read the result, fix, push again. That’s not a fast feedback loop. That’s the outer loop dressed up as a local development workflow.</p>
<p>And CI stayed green the entire time. The tests were passing. The test suite was technically healthy. The inner loop was broken.</p>
<p>The fix was migrating those test suites to <a href="https://blog.dicko.dev/posts/stop-copying-prod-into-dev-test-data-strategies-that-actually-scale/">Testcontainers</a> — self-contained, IDE-integrated, fast enough to run as part of the normal “Run All Tests” action. No terminal switching, no script execution, no timing-dependent startup. After the migration, local execution rates on those suites went from around 20% to 70% on days with active contributors. The full migration story — Docker-in-Docker CI gotchas, parallel execution, data seeding — is in <a href="https://blog.dicko.dev/posts/stop-copying-prod-into-dev-test-data-strategies-that-actually-scale/"><em>Stop Copying Prod Into Dev</em></a> if you want the technical detail. What matters here is that we didn’t find the problem through a retro, or an engineer complaint, or an architectural review. We found it because we were looking at a number nobody had been looking at before.</p>
<figure><img src="https://blog.dicko.dev/posts/what-we-actually-instrumented-and-what-we-found/image_2.webp" alt="" loading="lazy"></figure>
<h2>What the Coverage Actually Looks Like</h2>
<p>Across the .NET, JavaScript/TypeScript, and JVM stacks, the instrumentation now covers:</p>
<p>For .NET: compile time (incremental and full, via MSBuild), ASP.NET startup time, time to first response, and test run duration for both NUnit and xUnit.</p>
<p>For JavaScript and TypeScript: webpack and Vite full build time, HMR time (the one CI has never seen), and Jest and Vitest test run duration. Which we used for <a href="https://blog.dicko.dev/posts/the-40-minute-pipeline-why-your-build-system-is-making-you-dumber/">this comparison</a> between vite and rspack.</p>
<p>For JVM: Gradle via the Talaiot plugin, JUnit, and ScalaTest. JVM is the least mature coverage area — the collection libraries exist and are open source, but they haven’t received the same investment as the .NET and JS clients. If you work primarily in Kotlin, Scala, or Java and want to improve this, contributions are genuinely open and genuinely appreciated. The JVM ecosystem deserves the same quality of inner loop telemetry as .NET and JS.</p>
<p>The whole system is running across roughly 700 engineers at Agoda. That’s enough data to surface patterns that wouldn’t be visible at smaller scale — the hardware correlation being the clearest example — but the instrumentation is designed to be useful at team scale too.</p>
<h2>What Good Ownership Looks Like Here</h2>
<p>One thing worth naming explicitly: the dashboard doesn’t tell teams what to do with the data. It makes the data visible. What teams do with it is their call.</p>
<p>This is the same model as production monitoring. You wouldn’t centralise incident response into a platform team and have product teams ignore their own Grafana dashboards. The inner loop deserves the same ownership structure: the team that built the service is the team that runs it, and the team that runs it is the team that should own what their developer experience actually looks like. The platform provides the tooling. Teams own the signal.</p>
<p>The teams that have engaged with this most have done so because the data gave them something concrete to point at. Not “our builds feel slow” — a feeling, easy to dismiss. But “our first response time is 4 minutes 47 seconds at P75, and three months ago it was 40 seconds” — a number with a history, easy to act on.</p>
<h2>The Tool That Makes This Possible</h2>
<p>Everything above — the compile and startup split, the hardware correlation, the local vs CI execution comparison — runs through a single open-source dashboard: Agoda.DevExTelemetry. It&#39;s the missing piece: somewhere for the data to go that makes it visible, aggregatable, and actionable across teams.</p>
<p>Post 3 covers the architecture, the ingest endpoints, the dashboard views, and exactly how to get your own data flowing: <em>Introducing Agoda.DevExTelemetry</em>.</p>
<p>Or if you want to go straight to the repos:</p>
<ul>
<li><strong>Dashboard:</strong> <a href="https://github.com/agoda-com/Local-Dev-Telemetry-Manager">https://github.com/agoda-com/Local-Dev-Telemetry-Manager</a>- <strong>.NET clients:</strong> <a href="https://github.com/agoda-com/dotnet-build-metrics">https://github.com/agoda-com/dotnet-build-metrics</a>- <strong>JS clients:</strong> <a href="https://github.com/agoda-com/devfeedback-js">https://github.com/agoda-com/devfeedback-js</a>- <strong>JVM clients:</strong> <a href="https://github.com/agoda-com/java-local-metrics">https://github.com/agoda-com/java-local-metrics</a></li>
</ul>
<h2>Related Reading</h2>
<ul>
<li><a href="https://blog.dicko.dev/posts/the-inner-loop-nobody-measures/"><em>The Inner Loop Nobody Measures</em> </a>— Post 1 in this series; the case for why this matters before the how- <em>Introducing Agoda.DevExTelemetry</em> — Post 3; the open-source dashboard and setup guide- <a href="https://blog.dicko.dev/posts/bridging-worlds-making-net-bff-and-reactvite-play-nice-in-development/"><em>Bridging Worlds: Making .NET BFF and React/Vite Play Nice in Development</em></a> — the BFF and Vite setup that the frontend telemetry instruments- <a href="https://blog.dicko.dev/posts/stop-copying-prod-into-dev-test-data-strategies-that-actually-scale/"><em>Stop Copying Prod Into Dev: Test Data Strategies That Actually Scale</em></a> — the full Testcontainers migration story; what the telemetry found, this post fixed- <a href="https://blog.dicko.dev/posts/semantic-monitoring-the-question-youre-not-asking-about-your-production-systems/"><em>Semantic Monitoring: The Question You’re Not Asking About Your Production Systems</em></a> — same observability philosophy applied to production; this series is the development-side complement- <a href="https://blog.dicko.dev/posts/starting-a-paved-path-with-net-templates/"><em>Starting a Paved Path with .NET Templates</em></a> — where the telemetry clients belong in the paved path from day one- <a href="https://blog.dicko.dev/posts/mesh-programming-where-visual-design-meets-synchronized-development/"><em>Mesh Programming: Where Visual Design Meets Synchronized Development</em></a> — why inner loop speed directly affects collaborative design-engineering workflows</li>
</ul>
]]></content:encoded>
  </item>
  <item>
    <title>The Inner Loop Nobody Measures</title>
    <link>https://blog.dicko.dev/posts/the-inner-loop-nobody-measures/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/the-inner-loop-nobody-measures/</guid>
    <pubDate>Mon, 23 Mar 2026 00:04:10 GMT</pubDate>
    <category>developer-experience</category>
    <category>software-engineering</category>
    <category>software-development</category>
    <category>data-driven</category>
    <description>Or: Why Your Production Monitoring Is Excellent and Your Developer Experience Is a Black Box</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/the-inner-loop-nobody-measures/cover.webp" alt=""></figure>
<p>The sound wasn’t there anymore. That’s what he noticed first.</p>
<p>It was 10:23 on a Wednesday morning, and Somchai had pressed debug in Visual Studio, and the keyboard had gone quiet. The kind of quiet where your fingers are still hovering over the keys, not quite ready to let go of the idea that something might happen quickly. The build progress bar crept forward. The fan on the laptop started working harder than the conversation he was about to have with the code. He opened Slack. Read a thread. Closed it. The browser still wasn’t up. He pulled his Thai iced tea closer, the condensation already pooling on the vinyl floor beside his desk, and stared at the loading indicator with the particular thousand-yard look of someone who has been here before, recently, and will be here again this afternoon.</p>
<p>Nobody filed a ticket about it. There was no incident. The dashboards on level 6 were green. Standup had been fine. And somewhere, quietly, every developer on the team had developed a private protocol — their own small rituals to fill the dead time between pressing a key and seeing a result.</p>
<p>Here’s the uncomfortable truth: you are measuring the wrong thing. Almost certainly. And the gap between what you’re measuring and what your engineers actually experience every day is not a rounding error. It’s where the real cost of developer productivity disappears.</p>
<h2>The Outer Loop Has Excellent Monitoring</h2>
<p>We have genuinely gotten good at measuring what ships. DORA metrics and similar — deployment frequency, lead time for changes, Build time in CI, change failure rate, mean time to recovery — are well understood, increasingly instrumented, and legitimately useful. If you’re not tracking them, start. They are a reasonable proxy for organisational health and a good place to begin the conversation about engineering effectiveness.</p>
<p>But here’s what DORA and the others usually measure: the outer loop. The cycle from code committed to code in production. That’s one part of an engineer’s day. In many organisations, it’s not even the biggest part.</p>
<p>The inner loop — write, build, test, see — is where engineers spend most of their working hours. It’s the tight feedback cycle that determines whether a developer is in flow or context-switching into oblivion. And in most organisations, it is a complete black box.</p>
<p>As W. Edwards Deming observed, “If you can’t describe what you are doing as a process, you don’t know what you’re doing.” We’ve described the deployment process in careful detail and built excellent tooling to measure it. We’ve left the development process — the hours of actual engineering work that precede every PR — almost entirely undescribed.</p>
<h2>A Metric That Was Right About the Wrong Thing</h2>
<p>We had compile metrics. We were genuinely proud of them. Around 30 seconds at P75 for a non-trivial .NET service — that’s a decent build, I didn&#39;t believe it, people wouldn&#39;t complain if it was that good. I sat down with Somchai and asked him how it felt.</p>
<p>“When you press debug in Visual Studio and wait for the browser to come up — does that usually take about 30 seconds?”</p>
<p>He laughed. Not a polite laugh. The kind of laugh that means the question was almost charmingly naïve.</p>
<p>“Nooo,” he said. “About five minutes.”</p>
<p>I stared at him.</p>
<p>“But the build is only thirty seconds.”</p>
<p>He shrugged in that particular way that meant: <em>yes, and?</em></p>
<p>He was right. The build was thirty seconds. The compile metric was accurate. It was also measuring one link in a much longer chain: <strong>compile → startup → prewarm → browser ready</strong>. The service was loading real data from QA servers in the data centre at startup — database warm-up, cache population, dependency initialisation. All of it legitimate. None of it visible in our build metrics. The compile finished in 30 seconds. The service wasn’t usable for another four and a half minutes.</p>
<p>Because nobody had measured that gap, nobody knew it existed. And nobody knew it existed for every developer, every morning, every debug cycle.</p>
<p>This is what makes inner loop blindness particularly insidious: the metric wasn’t wrong. It was just measuring the wrong question. We were asking “how long does it take to compile?” when the question engineers were silently answering every day was “how long before I can actually see my changes?”</p>
<h2>The Chain Nobody Draws</h2>
<p>The inner loop is not a single event. It’s a chain, and every link can degrade independently.</p>
<p>For a backend service change, the chain looks something like this:</p>
<p><strong>compile → startup → prewarm / first request → browser ready</strong></p>
<p>For a frontend change:</p>
<p><strong>save → HMR → browser update</strong></p>
<p>For a test run:</p>
<p><strong>trigger → fixture spin-up → execution → result</strong></p>
<p>Most engineering organisations measure exactly one of these links in exactly one environment: CI build time. That’s the outer edge of the chain, run in a clean, well-resourced environment, on hardware nobody actually develops on day-to-day.</p>
<p>When HMR degrades from 200ms to 3000ms over three months of dependency additions, no CI metric catches it. When test suites become too slow to run locally, engineers stop running them — and CI stays green while the inner loop silently breaks. When startup time drifts from 30 seconds to five minutes because of a web server prewarm routine added three sprints ago, nobody files a ticket because nobody has a number to compare against. And its never 30 seconds it 5 minutes in 1 day, its slowly, over time.</p>
<p>Degradation without a baseline is just <em>how it is now</em>. Engineers adapt. They develop workarounds. They open Slack. They make another Thai iced tea. And the aggregate cost of those micro-interruptions — across a team of 50, or 200, or 700 — compounds into something that looks, from the outside, like a velocity problem with no obvious cause.</p>
<p>John Wooden, who spent decades coaching elite athletes, put it plainly: “It’s the little details that are vital. Little things make big things happen.” He was talking about basketball. He might as well have been talking about the 2 minutes of context-switching that happens every time a developer’s debug cycle is longer than it needs to be.</p>
<h2>Why This Is Harder to See Than You Think</h2>
<p>Production monitoring exists because production problems have visible consequences. Something goes down, something slows down, customers complain, revenue drops. The feedback loop is tight and the motivation to instrument it is obvious.</p>
<p>Inner loop degradation has none of those properties. It’s diffuse. It accumulates slowly. The cost manifests as slightly slower PR submission rates, slightly more context switching, slightly less willingness to run the full test suite before pushing. No alert fires. No dashboard turns red. Engineers adapt, rationally, to the environment they’re in — and the environment gets quietly worse.</p>
<p>There’s a structural problem too. The team that owns the monitoring platform is not the team that feels the pain of a slow debug cycle. Platform teams are incentivised to measure what they can control. What they can control is CI infrastructure, deployment pipelines, production observability. What they often cannot see is what happens on the developer’s laptop between the moment they start writing code and the moment they open a PR.</p>
<p>The same failure mode shows up in web performance — a centralised team owns the metric but isn’t the team that feels the consequence. The fix is the same: make the data visible to the teams who live with it, and give them the tools to act on it. You built it, you run it, you should own what your inner loop actually looks like. The platform team is here to help you, but they aren’t going to do the work for you, that’s on you.</p>
<h2>The One That Got Away</h2>
<p>Here’s the version of this problem that should keep you up at night: you don’t know what you’re missing.</p>
<p>Somchai knew his debug cycle was slow. He’d adapted to it. It was just <em>how things were</em>. The five minutes weren’t unusual to him because there was no comparison — no data showing that the rest of the team was seeing something different, no baseline that said “this used to be faster.” The metric gap didn’t just hide the problem from us. It hid it from the people experiencing it, because without a number to point at, slow becomes normal.</p>
<p>This is the failure mode that’s hardest to recover from. Not the fire that sets off alarms, but the slow drift that everyone adjusts to until nobody remembers what fast felt like.</p>
<p>The inner loop is where engineers live. It’s where flow happens or doesn’t. It’s where the best engineers on your team either find their rhythm or spend their afternoon watching progress bars. And until you measure it — all of it, not just the link that’s easiest to instrument — you’re operating on anecdote and intuition while calling it engineering.</p>
<h2>What Comes Next</h2>
<p>So what would you actually measure? And more importantly — what would you find that you weren’t looking for?</p>
<p>In the next post, we get specific: what we actually instrumented across .NET, JavaScript, and JVM stacks, how the collection works with zero friction for engineers, and what the data surfaced that had nothing to do with code. One finding involved three engineers, a webpack build, and hardware we didn’t know needed replacing. Another involved a test suite that was technically green in CI and functionally abandoned in practice.</p>
<p><em>Post 2: What We Actually Instrumented — and What We Found</em> — coming next.</p>
<p>And if you want to skip straight to the tool that makes all of this possible, <em>Post 3: Introducing Agoda.DevExTelemetry</em> covers the open-source dashboard and exactly how to get your own data flowing.</p>
<h2>Further Reading</h2>
<p>If you’re building the case internally for investing in inner loop observability, these are worth having in your back pocket:</p>
<ul>
<li>The <a href="https://queue.acm.org/detail.cfm?id=3595878">DevEx Framework from the ACM Queue</a> — feedback loops, flow state, and cognitive load as the three pillars of developer experience- <a href="https://itrevolution.com/product/accelerate/"><em>Accelerate</em></a> (Forsgren, Humble, Kim) — the DORA research. Good starting point for the outer loop argument; the inner loop argument is what comes after- The companion posts this series builds on: <a href="https://blog.dicko.dev/posts/bridging-worlds-making-net-bff-and-reactvite-play-nice-in-development/"><em>Bridging Worlds: Making .NET BFF and React/Vite Play Nice in Development</em></a> — the development setup we ended up instrumenting; <a href="https://blog.dicko.dev/posts/semantic-monitoring-the-question-youre-not-asking-about-your-production-systems/"><em>Semantic Monitoring: The Question You’re Not Asking About Your Production Systems</em></a> — the same philosophy applied to production; <a href="https://blog.dicko.dev/posts/mesh-programming-where-visual-design-meets-synchronized-development/"><em>Mesh Programming: Where Visual Design Meets Synchronized Development</em></a> — why inner loop speed directly affects collaborative design-engineering workflows; <a href="https://blog.dicko.dev/posts/the-impact-of-paved-paths-and-embracing-the-future-of-development/"><em>The Impact of Paved Paths and Embracing the Future of Development</em></a> and <a href="https://blog.dicko.dev/posts/starting-a-paved-path-with-net-templates/"><em>Starting a Paved Path with .NET Templates</em></a> — where telemetry clients belong in the paved path from day one</li>
</ul>
]]></content:encoded>
  </item>
  <item>
    <title>What It Looks Like When It’s Working</title>
    <link>https://blog.dicko.dev/posts/what-it-looks-like-when-its-working/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/what-it-looks-like-when-its-working/</guid>
    <pubDate>Sun, 22 Mar 2026 09:48:29 GMT</pubDate>
    <category>leadership</category>
    <category>engineering-culture</category>
    <category>agile</category>
    <category>software-engineering</category>
    <category>software-development</category>
    <description>Or: The Engineer Who Brings You Problems You Didn’t Know You Had</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/what-it-looks-like-when-its-working/cover.webp" alt=""></figure>
<p>It was a Tuesday afternoon — the quiet part, after lunch, when the open floor settles into that particular rhythm of keyboards and Thai iced tea and the low hum of people who are actually thinking rather than performing thinking. Somchai appeared at the side of my desk the way engineers do when they have something real to say — laptop open, not waiting to be invited. He had his laptop. He had a Grafana dashboard. He had a slide with three numbers on it and a very specific ask.</p>
<p>He didn’t have a Jira ticket. He didn’t have a complaint. He had a proposal.</p>
<p>I’ve spent three posts describing a problem and its causes. This one is about what the other side looks like — not as a theoretical destination, but as a specific, recognisable moment that you’ll know when you see it, if you know what you’re looking for.</p>
<p>That moment is an engineer walking in with a proposal nobody asked for, built on data they gathered themselves, describing a problem the organisation didn’t know it had — and telling you how they’re going to fix it.</p>
<p>That’s Station 4. And the reason it’s worth writing a whole post about is that it doesn’t announce itself. It arrives looking like a regular Tuesday afternoon conversation, and if you don’t recognise it for what it is — and respond to it correctly — you’ll handle it like a routine request and miss the signal entirely.</p>
<h2>The VP Definition</h2>
<p>A VP I worked with was once asked what a principal engineer actually looks like. He didn’t reach for a job level framework or a competency matrix. He said, without pausing:</p>
<p><em>“He’s the guy who brings me the problems I didn’t know I had and tells me how he’s going to solve them.”</em></p>
<p>That one sentence contains everything. Not “executes assigned work exceptionally well” — though they do. Not “flags problems when asked” — though they do that too. Surfaces problems the organisation didn’t know to look for, and arrives with a solution already forming. Both halves matter equally. The first half is about system understanding deep enough to see what’s invisible from above. The second half is ownership: treating the invisible problem as yours to solve, not yours to report.</p>
<p>Most teams have people capable of getting there. Most of those people never do — not because of ability, but because nobody ever made the gradient visible or told them which direction to move. The first three posts in this series were about removing the obstacles. This one is about what you’re removing them for.</p>
<h2>What Station 4 Actually Requires</h2>
<p>The leap from Station 3 to Station 4 isn’t just more of the same initiative. It requires a specific skill that most engineers were never taught and most managers never explicitly offer to teach: how to make a case.</p>
<p>Not a complaint. Not a gut feeling with a slide deck. A case — with a diagnosis, a quantified impact, a proposed solution, and an honest estimate of the cost. The kind of argument that could convince someone who doesn’t share your intuition about the system, because it’s built on evidence rather than authority.</p>
<p>This skill doesn’t develop through osmosis. It develops through a specific conversation, in a one-on-one, that goes roughly like this: <em>“You’re at the point where you should be bringing me proposals for larger investments. Here’s what a good one looks like.”</em> Then you show them. Not the slide template — the reasoning. What a well-framed problem looks like. What evidence is actually evidence versus noise. How to express the value of a technical improvement in terms a business can act on.</p>
<p>The engineers who never receive that conversation either never make the leap, or make it messily — arriving with proposals that are too vague to approve, getting knocked back, and concluding that advocacy doesn’t work here. The skill wasn’t missing. The scaffolding was.</p>
<h2>The Metrics Question</h2>
<p>Somchai’s three numbers were deployment frequency, mean time to recovery, and a third metric his team had built themselves — a composite score tracking how often engineers had to work around a known issue in the service rather than through it. Call it friction rate, for want of a better name.</p>
<p>The first two came from the DORA framework — the research documented in <a href="https://itrevolution.com/product/accelerate/"><em>Accelerate</em></a> by Forsgren, Humble, and Kim, which established deployment frequency, lead time for changes, change failure rate, and time to restore service as the four metrics that most reliably predict both engineering performance and organisational outcomes. If you haven’t read it, it’s worth your time. The evidence base is unusually solid for this domain, and it gives teams a common language for talking about engineering health that doesn’t require everyone to already agree on what good looks like.</p>
<p>But the third metric — the one his team invented — is the one worth paying attention to. Because DORA is a starting point, not a destination. It’s scaffolding for building the habit of measurement: the discipline of diagnosis, quantification, and evidence-based argument. Mature teams eventually develop metrics that are genuinely theirs — calibrated to their specific system, their specific failure modes, the business outcomes they’re actually accountable for.</p>
<p>An engineer who learns to make a case using DORA metrics has learned the skill. An engineer who looks at their system, decides DORA doesn’t capture the thing that actually matters here, instruments something new, and uses that to make a case — that’s a principal engineer. The framework taught them to measure. The judgment is theirs.</p>
<p>The proposal format that works, in my experience, has four parts. What’s the problem, stated precisely enough that someone who doesn’t work in the system could understand it. What’s the evidence that it’s actually a problem and not just an annoyance. What’s the proposed solution and its cost, in honest terms — not a best case, a realistic range. And what does success look like, specifically enough to know when you’ve achieved it. That’s it. Four parts, one page, fifteen minutes.</p>
<p>Somchai’s had all four. He was asking for six weeks.</p>
<h2>How to Respond When It Happens</h2>
<p>This is the part that matters most and gets written about least.</p>
<p>When an engineer walks in with a Station 4 proposal — unsolicited, evidence-based, specific — the correct response is not to immediately approve or reject it. The correct response is to engage with it seriously, in the room, in a way that makes visible that this is exactly the behaviour you were hoping for.</p>
<p>That means asking questions about the evidence, not the premise. Pushing back on the solution if you have concerns, but treating the problem framing as sound until proven otherwise. Being honest about constraints — other priorities, timing, resourcing — without those constraints becoming a reason to dismiss the proposal itself. And, regardless of the outcome, naming what just happened: <em>“This is exactly the kind of thing I want you bringing me.”</em></p>
<p>That last piece is not optional. The engineer who brought a Station 4 proposal and had it handled like a routine request — processed, filed, responded to in three days via Teams — has received a calibration signal whether you intended to send one or not. The signal is: this behaviour is normal. It isn’t. It took years to develop and it’s fragile in ways that aren’t obvious until it’s gone.</p>
<p>If you approve the proposal: say why, and say it in a way that reinforces the quality of the case, not just the idea. If you decline or defer: be specific about what would make it approvable, so the engineer leaves with something to work toward rather than a door closed. Either way, the conversation itself is the reward — being taken seriously as someone whose judgment about the system is trusted enough to act on.</p>
<h2>Closing the Loop on School</h2>
<p>In the <a href="https://blog.dicko.dev/posts/why-your-engineers-are-still-waiting/">first post</a>, I argued that the waiting instinct comes from education — from twelve-plus years of a system that trained engineers to execute within defined parameters and never once asked them to define the parameters themselves. That School, in the broadest sense, produced people who are exceptionally good at completing assignments and genuinely unprepared for the moment when nobody is assigning them.</p>
<p>Ralph Waldo Emerson wrote that <em>“the chief want in life is someone who will make us do what we can.”</em> Not allow. Not permit. Make — by creating conditions that stretch people toward what they’re actually capable of, rather than just clearing the path and hoping they find their way there.</p>
<p>The ownership ladder is that. The process flex from <a href="https://blog.dicko.dev/posts/the-process-that-punishes-the-behaviour-youre-building/">Post 2</a> creates the space. The cultural conditions from <a href="https://blog.dicko.dev/posts/ask-for-forgiveness-not-permission/">Post 3</a> make it safe. This post is what you’re building toward: engineers who have internalised the gradient, understand where they are on it, and are actively moving.</p>
<p>You don’t manufacture that. But you can make it possible — by naming the stations, giving explicit permission at each transition, responding correctly when someone moves, and building the kind of environment where a Tuesday afternoon conversation with a Grafana dashboard and a very specific ask is the most normal thing in the world.</p>
<p>That’s what it looks like when it’s working. You’ll know it when you see it.</p>
<h2>The Bottom Line</h2>
<p>The series started with a reflex — an engineer reaching for a browser tab to write a Jira ticket for something that would take twenty minutes to fix. It ends with an engineer who walks in on a Tuesday with three numbers, a proposal, and the confidence that the problem they found is worth someone’s attention.</p>
<p>The distance between those two moments isn’t talent. It’s the accumulated effect of a process that has enough flex to absorb initiative, a culture where failure is a question rather than a verdict, a manager who gave explicit permission at every rung of the ladder, and enough repetition of all three that acting became the reflex instead of waiting.</p>
<p>None of it is complicated. Very little of it is easy.</p>
<p>Now, if you’ll excuse me — Somchai got his six weeks. I should go find out what he’s built.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Ask for Forgiveness, Not Permission</title>
    <link>https://blog.dicko.dev/posts/ask-for-forgiveness-not-permission/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/ask-for-forgiveness-not-permission/</guid>
    <pubDate>Sat, 21 Mar 2026 21:30:49 GMT</pubDate>
    <category>engineering-culture</category>
    <category>psychological-safety</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>engineering-leadership</category>
    <description>Or: The Four Words That Changed How an Entire Room Understood Failure</description>
    <content:encoded><![CDATA[<p><em>Beer and Servers Don&#39;t Mix — Microservices Workshop Series, Part 3 of 4</em></p>
<hr>
<p>The slides had been up for forty minutes. Three quarters of missed targets, laid out in the kind of clean charts that make bad news look almost academic — until you do the arithmetic, which everyone in that room had done, silently, before the presenter got to the third slide. Three quarters. Not a blip. Not an unlucky run. A pattern. The people sitting closest to the presenter had, without seeming to notice they were doing it, shifted their chairs a few centimetres further away — the universal body language of people who have decided, at a subconscious level, that proximity to this moment carries a cost. The air conditioning hummed. Nobody moved.</p>
<p>The presenter finished. The room went the kind of quiet where you become very aware of your own breathing. Every eye moved to the end of the table.</p>
<p>The CEO said: <em>&quot;What did we learn?&quot;</em></p>
<hr>
<p>I&#39;ve been teasing that story across two posts. It earns the wait, I think, because the thing that makes it significant isn&#39;t what it did for the presenter. It&#39;s what it did for everyone else in the room.</p>
<p>Every person sitting there was running the same risk calculation, and four words answered it for all of them at once. Not a policy. Not a values statement in a slide deck nobody reads after the all-hands. A live demonstration, in the highest-stakes moment available, of what failure actually costs at this company. They all left that room knowing. And because they knew, each one of them was — fractionally, measurably, but genuinely — more willing to try something that might not work.</p>
<p>That&#39;s the fix. Not a process change. Not a new metric. A repeated, observable pattern of what happens when things go wrong — and the manager behaviour that either builds that pattern or destroys it.</p>
<hr>
<h2>Why Punishment Travels Faster Than Praise</h2>
<p>In the <a href="https://medium.com/beer-and-servers-dont-mix">last post</a>, we looked at the structural traps — the sprint metrics, the misaligned product person, the process that classifies initiative as failure. Those are real and they need fixing. But you can fix all of them and still have a team full of engineers waiting for permission they don&#39;t need, if the cultural signal underneath the process is wrong.</p>
<p>The cultural signal is set almost entirely by what leaders do when things go badly. Not what they say in retrospectives. Not what&#39;s written in the engineering principles doc. What they actually do, in the room, when someone took a risk and it didn&#39;t work out.</p>
<p>Here&#39;s the uncomfortable mechanism: punishment is more contagious than reward. When someone gets publicly criticised for acting without asking — overruled in a meeting, asked to &quot;run it by me first&quot; next time, had their initiative quietly un-done — the lesson doesn&#39;t land only on that person. It lands on everyone who saw it happen. The engineer sitting next to the engineer who got corrected updates their model faster than anyone. They weren&#39;t even involved. They didn&#39;t need to be. The signal was broadcast.</p>
<p>One incident, observed, can teach an entire floor of engineers that initiative isn&#39;t safe. That&#39;s not an exaggeration — it&#39;s how organisational learning actually works. We&#39;re social animals calibrating behaviour based on what we watch happen to each other, and we&#39;re very good at it.</p>
<p>The inverse is also true, but it takes longer to propagate and requires more repetition to stick. Which is why the CEO story matters as much as it does. Three quarters is a lot of evidence. The response, when it came, landed with proportional force.</p>
<hr>
<h2>The Manager Who Breaks It Without Knowing</h2>
<p>Theodore Roosevelt put it plainly: <em>&quot;The best executive is the one who has sense enough to pick good men to do what he wants done, and self-restraint to keep from meddling with them while they do it.&quot;</em></p>
<p>Self-restraint is the part most managers underestimate, because the moment someone starts genuinely taking ownership is precisely the moment managerial anxiety spikes. An engineer stops a sprint without asking. A developer makes a call on an architecture decision you would have made differently. Someone ships a small change and tells you about it afterward instead of before.</p>
<p>The instinct — entirely understandable — is to re-establish the loop. &quot;Just keep me in the loop.&quot; &quot;Could you check with me before decisions like this?&quot; &quot;I want to make sure we&#39;re aligned.&quot;</p>
<p>Each of these is a veto in process clothing. And it registers that way, even when it isn&#39;t intended that way. The engineer who just acted on ownership instinct — who is, in fact, doing exactly what you said you wanted — has now received a clear signal: the thing I did created friction. Don&#39;t do it again without asking.</p>
<p>The particularly insidious version is the manager who says all the right things about ownership and autonomy, and then responds to every exercise of that autonomy with &quot;that&#39;s great, but next time loop me in earlier.&quot; Said once, it&#39;s feedback. Said consistently, it&#39;s a policy. And it&#39;s a policy that teaches engineers your words and your behaviour are different things — and that your behaviour is the one that counts.</p>
<p>If your best engineers are still waiting for permission they don&#39;t need, the honest question isn&#39;t what&#39;s wrong with them. It&#39;s when was the last time someone acted without asking, and what did your face do.</p>
<hr>
<h2>What &quot;Rewarding Initiative&quot; Actually Looks Like</h2>
<p>The CEO didn&#39;t congratulate the product person for failing three quarters running. He asked what was learned — which carries an implicit expectation: that the failure produced data, that the data is now owned, and that it will be used. The reward isn&#39;t the failure. It&#39;s being treated as someone who tried, learned, and is still in the room.</p>
<p>That distinction matters because &quot;celebrate failure&quot; has become the kind of advice that sounds meaningful and means almost nothing in practice. Nobody is celebrating failure. What they&#39;re doing — when they do it well — is separating the act of trying from the outcome, and making it observable that the act of trying is what the organisation values.</p>
<p>Concrete behaviours that build this pattern over time:</p>
<p>When someone acts without asking and it works: praise the action, not just the outcome. &quot;Good call to just go for it&quot; lands differently than &quot;great result.&quot; The first rewards the decision-making. The second rewards the luck.</p>
<p>When someone acts without asking and it fails: &quot;What did we learn?&quot; Not &quot;how did this happen?&quot; Not &quot;why didn&#39;t you check with me first?&quot; The question signals that the information matters more than the accountability, and that the person asking it is interested in the former.</p>
<p>When someone asks permission for something they clearly could have just done: push it back. &quot;Why are you asking me? You know this better than I do.&quot; This one feels counterintuitive — surely it&#39;s good that they checked? — but what you&#39;re actually doing is interrupting the waiting reflex at the exact moment it surfaces, and replacing it with a signal that says: your judgment is sufficient here. Use it.</p>
<p>When you disagree with a decision someone made: be precise about what you&#39;re disagreeing with. &quot;I would have decided differently, and here&#39;s why&quot; is a conversation. &quot;You should have checked with me first&quot; is a punishment. The first develops judgment. The second teaches caution.</p>
<p>Amy Edmondson&#39;s research on psychological safety — documented in <em><a href="https://www.wiley.com/en-us/The+Fearless+Organization%3A+Creating+Psychological+Safety+in+the+Workplace+for+Learning%2C+Innovation%2C+and+Growth-p-9781119477242">The Fearless Organization</a></em> — shows consistently that teams where members feel safe to take interpersonal risks outperform those that don&#39;t, across a remarkable range of contexts and industries. The finding isn&#39;t that safety makes people comfortable. It&#39;s that safety makes people willing to do hard things. That&#39;s a different claim, and a more useful one.</p>
<hr>
<h2>Ask for Forgiveness, Not Permission</h2>
<p>The old line is usually treated as a piece of mildly cheeky personal career advice. It&#39;s actually an organisational design principle.</p>
<p>Environments where every action requires prior approval are environments where action slows to the speed of the approval queue. When initiative carries personal risk — a correction, a raised eyebrow, a &quot;why didn&#39;t you loop me in&quot; — people learn to wait. Not because they&#39;re timid. Because they&#39;ve correctly read the incentive structure and made the rational choice.</p>
<p>The phrase &quot;ask for forgiveness, not permission&quot; is really describing a specific kind of organisational trust: the trust that says your judgment is presumed sound until proven otherwise, rather than presumed unsound until cleared. Most engineering organisations say they extend this trust. Most engineering organisations, if you watch the behaviour rather than listen to the words, do not.</p>
<p>Building it requires consistency more than anything else. One public demonstration of the CEO question — &quot;what did we learn?&quot; — shifts a room. But one contradicting incident, where someone tries something and gets visibly corrected for the trying rather than the outcome, can undo several of those shifts at once. The asymmetry is frustrating and it&#39;s real. Culture is built slowly and damaged quickly, and the damage is always done in the moments when it&#39;s hardest to get it right.</p>
<hr>
<h2>The Permission You&#39;re Not Giving Explicitly Enough</h2>
<p>Here&#39;s the piece most engineering leaders miss: the permission to act without asking often needs to be stated directly, not just implied by the absence of rules against it.</p>
<p>Engineers who&#39;ve spent years in environments where initiative was risky don&#39;t automatically update their model when the environment changes. They need to hear it said out loud, in a one-on-one, specifically: <em>&quot;You&#39;ve been flagging this problem for three weeks. You don&#39;t need to ask me to fix it. If you see something that needs doing and you&#39;re confident in the call — do it. Tell me about it afterward.&quot;</em></p>
<p>That conversation, followed by a manager who responds to the first exercise of that permission with something other than friction, is worth more than any amount of &quot;we&#39;re an ownership culture here&quot; messaging. It&#39;s the difference between a policy and a pattern. And patterns are what people actually navigate by.</p>
<p>The station transitions from the <a href="https://medium.com/beer-and-servers-dont-mix">first post</a> — from waiting to reacting, reacting to initiating — don&#39;t happen by osmosis. They happen because a manager named them explicitly, gave specific permission for the next step, and then responded to the first attempt in a way that made the second attempt feel safe.</p>
<p>This is the work. Not the process reform from the last post, as necessary as that is. Not the metric changes. This: being the kind of leader who, when the room is watching and things have gone badly, asks what we learned — and means it enough that everyone in the room leaves knowing the answer.</p>
<hr>
<p>In the final post, we look at what it looks like when all of this is working. Station 4 ownership — engineers who bring you problems you didn&#39;t know you had and arrive with solutions already forming. How to teach the proposal format. How to know when someone is ready for it. And the metrics that tell you the culture has actually changed, rather than just gotten better at describing itself.</p>
<hr>
<hr>
<p><strong>Tags:</strong> Software Engineering, Engineering Leadership, Software Development, Tech, Engineering Culture, Psychological Safety, Management</p>
<p><strong>LinkedIn:</strong> Culture isn&#39;t what you say when things go well. It&#39;s what you do when everyone&#39;s watching and things have gone badly. Four words can change a room.</p>
]]></content:encoded>
  </item>
  <item>
    <title>The Process That Punishes the Behaviour You’re Building</title>
    <link>https://blog.dicko.dev/posts/the-process-that-punishes-the-behaviour-youre-building/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/the-process-that-punishes-the-behaviour-youre-building/</guid>
    <pubDate>Fri, 20 Mar 2026 14:24:24 GMT</pubDate>
    <category>engineering-leadership</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>engineering-mangement</category>
    <category>agile</category>
    <description>Or: Why Your Engineers Stopped Taking Initiative — and Scrum Helped</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/the-process-that-punishes-the-behaviour-youre-building/cover.webp" alt=""></figure>
<p><em>Ownership Series, Part 2 of 4</em></p>
<figure><img src="https://blog.dicko.dev/posts/the-process-that-punishes-the-behaviour-youre-building/image_1.webp" alt="" loading="lazy"></figure>
<p>It was 11:20 on a Thursday morning, and the sprint review had been going for six minutes before the energy in the room shifted. Not dramatically — just the subtle recalibration that happens when everyone simultaneously realises where a conversation is heading. Namfon had pulled two days out of the sprint to fix a performance issue that had been silently degrading the page they’d spent the last three sprints building — interaction times creeping up, the kind of thing users feel before they can name it, the kind of thing that makes a Product Owner’s carefully planned feature land softer than it should have. She hadn’t asked. She’d seen it, understood it, and fixed it. The work was good. The fix was right. The page hadn’t dragged since.</p>
<p>The sprint, however, was two points short of its commitment. And the Jira burndown on the big monitor on the wall made sure everyone in the room knew it.</p>
<p>In the<a href="https://blog.dicko.dev/posts/the-process-that-punishes-the-behaviour-youre-building/"> last post</a>, we looked at the problem: engineers who default to waiting, who ticket small fixable problems instead of fixing them, who treat ownership like a school assignment rather than something they’ve genuinely claimed. We traced the instinct back to where it comes from — twelve-plus years of education that rewards execution within defined parameters and never once asks you to define the parameters yourself.</p>
<p>This post is about something harder to fix. Because the waiting instinct, at least, is just conditioning. You can interrupt conditioning with enough repetition and enough visible demonstration of the alternative.</p>
<p>What you can’t interrupt as easily is a process that’s actively working against you.</p>
<h2>The Scrum Trap</h2>
<p>Here’s the structural irony at the centre of most engineering organisations: the behaviour that defines Station 3 ownership — stopping to fix something important that nobody asked you to fix — is precisely the behaviour that most Scrum implementations are designed to punish.</p>
<p>Not intentionally. Not maliciously. But effectively.</p>
<p>Sprint completion rate is a standard health metric in virtually every Scrum team I’ve encountered. When an engineer decides mid-sprint to address unplanned work, the committed scope doesn’t get delivered. The burndown chart goes red. In many teams, the whole team is accountable for the miss — velocity drops, and someone in a management review asks why the team keeps not hitting its commitments. The message lands on everyone, not just the person who acted. One engineer does something. The whole platoon does push-ups.</p>
<p>As W. Edwards Deming observed: <em>“A bad system will beat a good person every time.”</em></p>
<p>Deming was talking about manufacturing — the assembly line worker blamed for defects that the production system itself was generating. But the mechanism is identical here. You can hire engineers with excellent instincts for ownership, put them in an environment where exercising those instincts makes their team miss its sprint, and watch those instincts get quietly trained out of them inside six months. It’s not a character failure. It’s a rational adaptation to the incentive structure they’re operating in.</p>
<p>Namfon did the right thing. The system told her she didn’t.</p>
<h2>What the Manifesto Actually Says</h2>
<p>Here’s the part that should make every certified Scrum Master slightly uncomfortable.</p>
<p>The Agile Manifesto — the document that Scrum claims as its foundational text — contains twelve principles, not four values. And principle number nine reads: <em>“Continuous attention to technical excellence and good design enhances agility.”</em></p>
<p>Not “when the backlog allows for it.” Not “subject to Product Owner prioritisation.” Continuous attention.</p>
<p>The manifesto’s authors understood that agility and technical health aren’t in tension — they’re the same thing. A team that stops to fix something important isn’t deviating from the spirit of Agile. They’re living it. The sprint commitment model, applied rigidly, has managed to produce the opposite of what the document intended: a process where the most agile response to a discovered problem — fixing it immediately, before it gets worse — is classified as a sprint failure.</p>
<p>We took a manifesto that explicitly values responding to change over following a plan, and built an industry around following the plan.</p>
<h2>The Velocity Illusion</h2>
<p>The deeper problem with sprint completion metrics is what they measure and what they don’t.</p>
<p>Velocity measures output. It counts points delivered, features shipped, tasks completed. It has nothing to say about the quality of what was delivered, the health of the system it was delivered into, or the compounding cost of the things that were deferred to keep the number looking good.</p>
<p>A team that ships every sprint on schedule while ignoring a quietly growing reliability problem has excellent velocity and deteriorating capacity. Those two facts don’t appear on the same dashboard. By the time the capacity deterioration becomes visible — slower delivery, more bugs, longer debugging sessions, engineers spending half their time working around problems rather than solving them — the connection to the sprint metrics that enabled it has become almost impossible to make.</p>
<p>John Wooden, who won ten NCAA basketball championships at UCLA, used to tell his players: <em>“Never mistake activity for achievement.”</em> He wasn’t talking about sprint planning, but he might as well have been. A team that’s busy isn’t necessarily a team that’s improving. And a team that occasionally pauses to fix something important is doing more for its long-term velocity than any number of fully-committed sprints.</p>
<p>The teams that understand this shift their definition of a successful sprint away from scope delivered and toward outcomes achieved. Did the system get more reliable? Did the deployment get faster? Did the thing we were worried about last quarter get less worrying? Those questions have answers too — they just require different metrics, and a willingness to accept that some of the most valuable work a sprint can contain is work that wasn’t on the board when it started.</p>
<h2>The Product Alignment Problem</h2>
<p>Even if you fix the process — build in the slack, shift the metrics, create explicit permission for mid-sprint course correction — there’s a second structural trap waiting.</p>
<p>The misaligned product person.</p>
<p>An engineer starts taking initiative. They fix things without being asked. They stop a sprint to address something they own. Their engineering manager is giving them positive signals — “good call, that needed doing.” Then a PM pushes back. “Why wasn’t this in the sprint plan?” “We agreed on scope.” “I need to be able to predict what this team delivers.”</p>
<p>The engineer is now caught between two authority figures sending directly contradictory signals about what good looks like. No single person is the villain. The engineering manager is correctly rewarding ownership. Product is correctly managing predictability. The engineer — who was just making the leap from Station 2 to Station 3 — quietly concludes that initiative is conditionally safe at best, and goes back to waiting.</p>
<p>This failure mode is more common than the manager who actively punishes initiative, because it doesn’t feel like punishment from either side. It feels like a process disagreement. But to the engineer watching it play out, the signal is identical: acting without clearance has a cost.</p>
<p>The fix requires a conversation that happens <em>before</em> an engineer gets caught in the crossfire. Engineering and product need to be aligned on a few non-negotiable principles: sub-day fixes are never sprint failures, they’re engineering hygiene; an engineer stopping to fix something they own is categorically different from scope creep; and the goal of sprint planning is outcomes, not exhaustive pre-commitment of every available hour. Without that alignment, you’re asking engineers to take initiative inside a system that will occasionally, unpredictably, penalise them for it. That’s not ownership culture. That’s a minefield.</p>
<h2>The Flex You Need to Build</h2>
<p>None of this means abandoning sprint planning or throwing predictability out the window. It means building enough flex into the process to absorb initiative without classifying it as failure.</p>
<p>In practice, this looks like a few specific decisions:</p>
<p><strong>Explicit capacity slack.</strong> Some teams reserve a percentage of sprint capacity — call it ten or fifteen percent — that’s understood to be available for unplanned good work. Not technical debt in the abstract, not “innovation time” that mysteriously vanishes when commitments get tight, but a deliberate buffer that signals: we expect to find things worth fixing, and we’ve made room for that.</p>
<p><strong>A team norm about quality stops.</strong> The explicit, stated understanding that stopping to fix something important is never a sprint failure — it’s a data point. The debrief question isn’t “why didn’t you finish the committed scope?” It’s “what did we find, and was fixing it the right call?” Usually the answer is yes. Occasionally it’s a useful conversation about prioritisation. Either way, nobody gets push-ups.</p>
<p><strong>Outcome metrics alongside output metrics.</strong> Deployment frequency. Lead time for changes. The number of times the team had to work around a known problem this sprint rather than through it. These sit alongside velocity on the dashboard rather than replacing it — and they make visible the things that velocity alone is hiding.</p>
<p>The goal is a process that can absorb a Namfon moment — an engineer who sees something important and fixes it before asking — without the burndown chart becoming an accusation.</p>
<h2>What’s Actually in the Way</h2>
<p>So if you’re an engineering leader reading this and recognising your own team in the waiting pattern from the last post, the honest question isn’t “how do I change my engineers?” It’s “what is the process telling them?”</p>
<p>Because your engineers are not ignoring their instincts. They’re following them. They’ve learned, through observation, that the environment rewards completion of assigned work and occasionally penalises deviation from it. They’ve watched what happens when someone acts without clearing it first. They’ve felt the implicit weight of a sprint that went red because someone did the right thing at the wrong time.</p>
<p>The instinct to wait isn’t irrational. In most Scrum environments, it’s the correct adaptation.</p>
<p>That’s the trap. And it’s one that no amount of “we want engineers who take ownership” messaging will fix, because the message and the system are saying different things — and the system always wins.</p>
<p>In the next post, we get to the fix. There’s a specific story about a CEO in a room full of people watching him respond to three consecutive quarters of failure, and it’s the clearest illustration I’ve found of how culture actually changes — not through what leaders say when things go well, but through what they do when everyone’s watching and things have gone badly.</p>
<p>The fix, it turns out, isn’t a process change. It’s four words.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Why Your Engineers Are Still Waiting</title>
    <link>https://blog.dicko.dev/posts/why-your-engineers-are-still-waiting/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/why-your-engineers-are-still-waiting/</guid>
    <pubDate>Fri, 20 Mar 2026 14:03:30 GMT</pubDate>
    <category>engineering-culture</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>agile</category>
    <category>engineering-leadership</category>
    <description>Or: The Ownership Problem Nobody Admits They Created</description>
    <content:encoded><![CDATA[<p><em>Beer and Servers Don&#39;t Mix — Microservices Workshop Series, Part 1 of 4</em></p>
<hr>
<p>He&#39;d been staring at it for eleven minutes. You could tell by the way his cursor had stopped moving — hovering somewhere in the middle of the screen, that specific stillness that isn&#39;t thinking, it&#39;s deciding whether to think. The standup had finished, the vinyl floor outside still humming with the footsteps of people heading back to their desks, and Somchai was looking at a bug that was going to take him about forty minutes to fix. He knew it. The team knew it. Everyone had just agreed, out loud, that it was the kind of thing that should probably get done soon.</p>
<p>He opened a browser tab and started writing a Jira ticket.</p>
<hr>
<p>You&#39;ve seen this. You may have done this. The instinct to ticket a small, fixable problem rather than fix it is so deeply embedded in most engineering teams that it doesn&#39;t even register as a decision anymore — it just happens, like reaching for your phone when a conversation goes quiet. And if you&#39;re an engineering leader watching it happen, the right response probably isn&#39;t frustration. It&#39;s curiosity. Because that reflex didn&#39;t come from nowhere, and it didn&#39;t come from laziness.</p>
<p>It came from twelve-plus years of school.</p>
<hr>
<h2>The Education You Didn&#39;t Know You Were Getting</h2>
<p>Here&#39;s the thing nobody puts in a job description: most engineers arrive at their first role exceptionally well-trained to wait for instructions.</p>
<p>Not because they&#39;re passive. Not because they lack ambition. Because school — from the first day of primary school to the last exam of university — is a long, carefully structured sequence of assigned work assessed by authority figures. Every task has an owner, and that owner isn&#39;t you. Your job is to complete what&#39;s been defined, meet the deadline that&#39;s been set, and wait for the next assignment. The system is optimised to produce people who execute well within defined parameters. It is not optimised to produce people who walk up to an unsolved problem and say &quot;that&#39;s mine.&quot;</p>
<p>Then those people join your team. You hand them a system to own. And for the first several months — sometimes longer — they treat it exactly like a school assignment. They do what&#39;s asked of them. They do it well. They&#39;re responsive, technically sharp, diligent. But they wouldn&#39;t dream of stopping a sprint to fix something nobody told them to fix. The system is assigned to them. They haven&#39;t made the leap to <em>owning</em> it.</p>
<p>Those are different things. And the gap between them is almost entirely invisible until you know what to look for.</p>
<hr>
<h2>What &quot;Waiting&quot; Actually Looks Like</h2>
<p>The Jira ticket is the tell, but it&#39;s not the only one.</p>
<p>Waiting looks like a standup where every update is about tasks on the board — never about things adjacent to the board that are quietly on fire. It looks like an engineer who flags a problem in a one-on-one but wouldn&#39;t raise it in planning, because planning is where the Product Owner decides what matters. It looks like a pull request that fixes exactly what was asked, no more, even when the author clearly noticed three related things that would have taken ten minutes each to address.</p>
<p>None of this is incompetence. It&#39;s calibration — the kind of calibration that happens when people learn, through observation, what the environment actually rewards. And most engineering environments, without meaning to, reward completion of assigned work over initiative. They reward staying in your lane. They reward asking before acting.</p>
<p>The economist Milton Friedman once observed that <em>&quot;one of the great mistakes is to judge policies and programs by their intentions rather than their results.&quot;</em> He was talking about government, but the mechanism is identical in engineering organisations. You can intend to build an ownership culture. You can say &quot;we want engineers who treat their systems as genuinely theirs.&quot; And then you can build, entirely by accident, an environment that teaches them the opposite — and wonder why nothing changes.</p>
<hr>
<h2>The Spectrum Nobody Tells You About</h2>
<p>Ownership isn&#39;t a switch you flip. It&#39;s a gradient, and most engineers spend their entire careers at one end of it without ever being told there&#39;s another end to move toward.</p>
<p>The gradient has four recognisable stations — and they matter because the manager&#39;s job is genuinely different at each one.</p>
<p><strong>Station 1: Waiting.</strong> Executes what&#39;s asked. Good work, no initiative. This is where almost everyone starts, and it looks fine until you realise nothing is happening that you didn&#39;t specifically request.</p>
<p><strong>Station 2: Reacting.</strong> Flags problems they notice. Raises things in standups that weren&#39;t on the agenda. Advocates — once, maybe twice — for fixing something that&#39;s bothering them. This is the first real signal that something is shifting. They&#39;re paying attention to the health of the system, not just the tasks on the board.</p>
<p><strong>Station 3: Initiating.</strong> Stops a sprint for a couple of days to fix the thing they&#39;ve been flagging for three weeks. Makes a call and tells you about it afterward. Doesn&#39;t ask permission for every small decision because they&#39;ve worked out they don&#39;t need to. This is where &quot;assigned responsibility&quot; actually becomes ownership — and, as we&#39;ll explore in the next post, it&#39;s also where most organisational systems inadvertently punish them for it.</p>
<p><strong>Station 4: Advocating.</strong> Comes to you with a proposal. Not a complaint dressed as a suggestion — a real case. They&#39;ve looked at the data, framed the problem in terms the business can act on, estimated the impact, and asked for two sprints of investment. This is principal engineer behaviour. It requires seeing problems the organisation didn&#39;t know to look for, and treating those problems as yours to solve.</p>
<p>A VP I worked with once described a principal engineer in one sentence. <em>&quot;He&#39;s the guy who brings me the problems I didn&#39;t know I had and tells me how he&#39;s going to solve them.&quot;</em></p>
<p>That&#39;s Station 4. Most teams have a handful of people capable of getting there. Most never do — not because of ability, but because nobody ever made the gradient visible, let alone told them they were supposed to be climbing it.</p>
<hr>
<h2>Back to the Ticket</h2>
<p>So. Somchai&#39;s ticket.</p>
<p>My response, when I see this happen, is to fix it myself. Open a pull request. Then make the time visible.</p>
<p>The critical detail is the timestamp. When I paste the PR link in Slack — sometimes less than twenty minutes after the Jira tab opened — I&#39;m not making a point about Somchai. I&#39;m collapsing an assumption the whole team shares: that fixing it yourself is the slow path, the risky path, the path that requires clearance. The time to write the ticket and the time to fix the thing are roughly the same. Often less. The ticket isn&#39;t efficiency — it&#39;s the appearance of efficiency, wrapped around an instinct to defer.</p>
<p>Then I show up at standup. Say it out loud. Not as a reprimand — as a demonstration. &quot;This isn&#39;t a Jira ticket. It&#39;s a twenty-minute fix. The ticket takes longer to write than the problem takes to solve.&quot; I reply in the Slack thread with the same message. I make the pattern visible and name it as a pattern, because that&#39;s the only way to interrupt an automatic response: show the alternative, repeatedly, until the alternative becomes the reflex.</p>
<p>It takes longer than you&#39;d think. The instinct runs deep.</p>
<hr>
<h2>What This Series Is About</h2>
<p>This is the first of four posts on the ownership ladder — how engineers move from waiting to initiating to advocating, and what actually gets in the way.</p>
<p>Post 1 (this one) is about the wound: what waiting looks like, where it comes from, and why handing someone ownership rarely produces the ownership you were hoping for.</p>
<p>Post 2 is about the trap: the structural reasons — Scrum, sprint metrics, product alignment — that punish Station 3 behaviour before it even has a chance to become a habit.</p>
<p>Post 3 is about the fix: the cultural conditions, and the specific manager behaviours, that make taking ownership safe. There&#39;s a story in that post about a CEO in a room full of people watching him respond to three consecutive quarters of failure. It&#39;s worth the wait.</p>
<p>Post 4 is about the proof: what it looks like when it&#39;s working, how to measure it, and the specific conversation you need to have when someone is ready to move from Station 3 to Station 4.</p>
<p>For now: next time someone on your team tells you they&#39;re going to create a Jira ticket for something — check your watch. Then open your IDE.</p>
]]></content:encoded>
  </item>
  <item>
    <title>The Thematic Quarter: Why Your OKRs Are Fighting Each Other</title>
    <link>https://blog.dicko.dev/posts/the-thematic-quarter-why-your-okrs-are-fighting-each-other/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/the-thematic-quarter-why-your-okrs-are-fighting-each-other/</guid>
    <pubDate>Sun, 15 Mar 2026 07:35:28 GMT</pubDate>
    <category>engineering-leadership</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>okr</category>
    <category>agile</category>
    <description>Or: How a Book About Silos Changed How We Set Goals</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/the-thematic-quarter-why-your-okrs-are-fighting-each-other/cover.webp" alt=""></figure>
<p>It was a Thursday morning on level 6 and the glass-walled meeting room was full in the way that only quarterly planning meetings are full — every chair taken, a few people standing along the back wall with their arms folded the same way, cups of Thai iced tea going warm on the table because nobody wanted to be the one making noise with a straw. The slide on the screen had forty-three goals across nine teams, colour-coded into something that must have looked like alignment when someone built it at midnight the night before. Namfon was nodding slowly in the second row — the kind of nod that means nothing, the professional nod, the nod that says <em>I am present and processing</em>. Through the glass, you could see the rest of the floor — heads down, headphones on, entirely indifferent to whatever was being decided in there.</p>
<p>We’ve all been in that room. Some of us have <em>run</em> that room. And the problem isn’t the colour-coding.</p>
<h2>The OKR Promise That Doesn’t Quite Deliver</h2>
<p>OKRs are a genuinely good idea. John Doerr <a href="https://www.whatmatters.com/">wrote the book on them</a> — quite literally — and the framework is sound: set ambitious goals, measure outcomes, build accountability. The problem isn’t the framework. The problem is the assumption buried underneath it.</p>
<p>That assumption is this: if every team sets good goals, the organisation will have good goals.</p>
<p>It won’t. And if you’ve led more than two or three teams, you already know this in your bones even if nobody’s said it plainly. You can have ten teams all hitting their OKRs — genuinely, not just cooking the numbers — and still end the quarter with the organisation having made no meaningful progress on the thing that actually matters. Because each team optimised locally. Each team did exactly what it was measured on. And local optimisation, done well, doesn’t automatically sum to global progress.</p>
<p>Fredmund Malik, the Austrian management theorist, put it plainly: <em>“The greatest enemy of organisational effectiveness is not incompetence. It is local optimisation.”</em> He wasn’t being clever. He was describing the thermodynamics of how organisations actually work: energy put into one part of the system doesn’t automatically flow to where it’s needed most. It stays where it was generated.</p>
<p>This is the OKR problem nobody wants to say out loud. Teams have goals. Teams pursue them. The organisation drifts.</p>
<h2>Enter Lencioni — and the Thematic Goal</h2>
<p>Patrick Lencioni’s <a href="https://www.tablegroup.com/books/silos-politics-and-turf-wars/"><em>Silos, Politics and Turf Wars</em></a> is a book most engineers won’t read because it’s written as a business fable and it doesn’t mention a single line of code. That’s a shame, because the central insight is one of the most practically useful things I’ve encountered in over two decades of leading engineering teams.</p>
<p>Lencioni’s argument is deceptively simple: silos aren’t caused by bad people, territorial managers, or bad culture. They’re caused by the <em>absence of a compelling shared objective</em> — one single goal that temporarily overrides everything else, one thing that everyone in the room is trying to accomplish together. He calls this a <strong>thematic goal</strong>.</p>
<p>Not a set of OKRs. Not a strategy doc. Not a vision statement. One goal. Qualitative, not quantitative. Defined for a specific period. Shared across the whole leadership layer in a way that’s specific enough to hurt.</p>
<p>Lencioni arrives at this idea through an observation that’s worth understanding: working across several companies, he noticed that one of them didn’t seem to have a silo problem when all the others did. The reason turned out to be a crisis the company was already living through — the kind where if they didn’t solve it, everyone goes home. That level of stakes had pulled the organisation together in a way that normal operations never had. Everyone knew what mattered. Everyone knew why. Nobody was protecting territory because there was no territory worth protecting if the whole thing collapsed. The crisis had manufactured the shared context that made collaboration the obvious default rather than the effortful exception.</p>
<p>His insight — and it’s a sharp one — is that crises create exactly the shared context that makes people collaborate. The bad news is that most organisations don’t figure this out until they’ve nearly destroyed themselves. The thematic goal is Lencioni’s answer: a deliberately manufactured shared context that mimics the clarity of a crisis without requiring the near-death experience. You want the urgency. You don’t want the stakes.</p>
<h2>I’ve Seen the Crisis Version</h2>
<p>Before Agoda, I worked at a ticketing company based in Australia, doing ticketing for large festival seasons. Once a year, the big onsales hit: hundreds of thousands of customers simultaneously trying to buy tickets for the same events. The system didn’t always cope.</p>
<p>On one occasion, we took a telephone exchange offline. Not metaphorically. The call volume from customers trying to buy tickets overwhelmed the local network infrastructure. An entire telephone exchange. Offline.</p>
<p>In that moment, nobody in the office asked what the quarterly priorities were. Nobody checked the backlog. Nobody talked about process. Every single person in that building knew exactly what needed to happen and who needed to talk to whom to make it happen. The team in Bangkok and the team in Australia, who might otherwise have traded polite emails for days, were on the phone with each other continuously. The onsale season, every year, was a voluntary manufactured crisis — a period where everyone dropped territorial behaviour because the shared problem was too obvious and too immediate to ignore.</p>
<p>That’s Lencioni’s thematic goal at its most visceral. You don’t want to <em>need</em> a crisis to feel that way. But you do want to find a way to bottle it.</p>
<p>Jim Collins, in <a href="https://www.jimcollins.com/books/g2g.html"><em>Good to Great</em></a>, wrote something that applies here at a different level than most people use it: <em>“If you have more than three priorities, you don’t have any.”</em> That’s usually quoted in the context of individual or team focus, but the more important application is <em>across</em> teams. Across ten teams, you have fifty priorities. None of them are shared. Collins is making an argument about concentration of force. The thematic goal is how you apply that force at the organisational level.</p>
<h2>Breaking the Monolith: What a Thematic Quarter Actually Looks Like</h2>
<p>Let me give you a concrete example from Agoda, because abstract frameworks are only as useful as the stories you can attach to them.</p>
<p>A few years ago, across roughly ten engineering teams, we ran what I’d now describe as a thematic year — though we wouldn’t have used that term at the time. The goal was structural: we had a large central monolith that one team owned, and that ownership had made them a bottleneck for every other team that needed to ship anything touching their domain. The solution was decomposition — breaking the monolith into bounded services, each owned by the product team closest to that business domain.</p>
<p>In principle, this was a clean win. In practice, it required nine teams that had previously been <em>consumers</em> of another team’s work to become <em>contributors</em> to a shared codebase they’d never owned. Cross-team code reviews. Shared responsibility for deployment pipelines. Design decisions that couldn’t be made unilaterally. That friction was the shared investment that made the goal real — this wasn’t one team’s project that everyone else watched from the sidelines. Every team gave something up.</p>
<p>A few things made it work that wouldn’t have been obvious in advance.</p>
<p><strong>A specialist team took the first heavy lifting.</strong> A dedicated group built the initial project templates, the shared CI infrastructure, and did enough of the first migration work that product teams had a pattern to follow. The activation energy problem is real: people won’t change how they work if the new way of working requires them to invent everything from scratch while also delivering a quarter. Give them the first move, and they’ll follow.</p>
<p><strong>Measurement was on the right metric — and it wasn’t obvious.</strong> We tracked merge requests into the monolith versus the new services, and we deliberately prioritised moving the components that had the <em>highest contribution volume first</em> — the parts people were most actively changing, not the parts that seemed most architecturally central. This is counterintuitive. You might think you should start with the core. But starting with what’s most actively touched means you immediately reduce the friction that teams feel every day. Progress becomes visible to the people who are doing the work.</p>
<p><strong>The goal was the question in every 1:1.</strong> Not occasionally. Not quarterly. Every relevant technical review, every planning session, every OKR check-in: <em>“How is this helping with the monolith split? If it’s not, why are you doing it?”</em> Not as a gotcha. As a genuine coordination question. (This is the mechanism that makes culture stick — I wrote about it in the context of <a href="https://blog.dicko.dev/posts/the-signature-question-how-culture-actually-spreads/">building repeating leadership habits</a> — the same principle applies at the goal level. If the goal isn’t the question you ask in every room, it isn’t the goal.)</p>
<p><strong>A monthly cross-team demo kept progress social.</strong> With ten teams, a monthly demo where everyone sees what everyone else is shipping is tractable. The monolith progress was visible and accumulating. Teams could see each other’s work. The goal had a face. Past much beyond ten teams this stops working — the demo becomes too long, too broad, and too passive to be useful. If you’re operating at that scale you’ll need a different mechanism for cross-team visibility, but the underlying need is the same: progress has to be seen, not just reported.</p>
<p><a href="https://blog.dicko.dev/posts/recognition-the-leadership-skill-nobody-taught-you/"><strong>Recognition</strong></a> <strong>made the goal visible beyond the room.</strong> I had t-shirts made — “Monolith Breaker” — and gave them out to engineers who did meaningful work toward the goal. Not as a formal award programme. As a moment: a photo of the presentation, posted in a large public Slack channel, visible to the whole engineering floor. It sounds small. It isn’t. When people see a colleague being recognised publicly for contributing to a specific goal, two things happen: the goal becomes real to everyone watching, and the recognised engineer becomes a visible ambassador for the work. The t-shirt is a prop. The Slack post is the message. Both together are a reminder, repeated across the year, of what the organisation is actually trying to do.</p>
<figure><img src="https://blog.dicko.dev/posts/the-thematic-quarter-why-your-okrs-are-fighting-each-other/image_2.webp" alt="" loading="lazy"></figure>
<p>Three years in, we’re in the tail end. Not finished — these things always have a long tail — but past the heavy phase. After the first year, 80% of pull requests were going into the new services and only 20% were still going into the monolith. The code volume was probably the inverse ratio — there was still a lot of monolith left — but that’s the point. We weren’t trying to move the most code first. We were trying to move the most <em>activity</em> first, the parts people were touching every day. The impact was felt long before the migration was complete. The first year moved the most. That’s typical: the first push creates momentum, the middle phase grinds, the tail is long. If you’re planning a thematic year that involves structural change, build your expectations around that arc.</p>
<h2>Why “Shared” Goals Usually Aren’t</h2>
<p>The failure mode I’ve watched play out — in different organisations, at different scales — tends to look like one of two things.</p>
<p>The first is the <strong>waterfall of OKRs</strong>. The goal gets set at senior level, translated into team OKRs, and by the time it reaches engineers it’s unrecognisable. Everyone is technically “contributing to the goal” in a way that’s defensible and meaningless. The goal has been waterfall-planned to death before any work starts. You define the goal at the top, then let the waterfall of translation drain all the meaning out of it before it reaches anyone who can act on it.</p>
<p>The second is the <strong>announcement and retreat</strong>. The goal is announced. Everyone nods. Teams return to their existing work, attach a thin justification layer to their existing plans, and proceed as before. The thematic goal becomes a filter on how people <em>describe</em> what they were going to do anyway — not a driver of what they actually do.</p>
<p>Both failure modes have the same root cause: <strong>there wasn’t much effort on actually pushing the goal down.</strong> The announcement happened. The OKRs were set. Leadership assumed the goal had gravity of its own.</p>
<p>It doesn’t. Not at scale.</p>
<p>At ten teams, the informal cross-team connections needed for coordination exist or can be created with light effort. The monthly demo is tractable. One director can maintain real visibility into all ten teams. The mechanism works with relatively light management overhead. This broadly maps to what <a href="https://less.works/">Large-Scale Scrum (LeSS)</a> recommends: roughly four to eight teams per coordination area before you need to break into separate coordination structures.</p>
<p>Past around ten teams, the mechanism needs deliberate reinforcement that most organisations don’t provide. The informal connections aren’t there to be activated. The goal can’t travel through them. The OKRs are set and then left to fend for themselves. I’ve watched the failure play out at 22 teams — the goal gets announced, the OKRs cascade, and nine months later every team has a coherent story about how they contributed to the goal and the goal hasn’t moved. The stories are all true. The goal is still in the same place.</p>
<h2>Two Things a Thematic Goal Actually Needs</h2>
<p>The standard OKR advice — be measurable, be ambitious, be time-bound — is all sound. It’s also insufficient. A thematic goal needs two things that the standard advice doesn’t address.</p>
<h2>1. Shared Pain — and a Shared Problem</h2>
<p>A goal that only inconveniences one team isn’t thematic. It’s that team’s priority with better branding.</p>
<p>The test is honest and uncomfortable: can every team in the room articulate what <em>they</em> had to deprioritise or change because of this goal? If the answer is “nothing, really,” you don’t have a thematic goal. You have one team doing hard work while everyone else offers moral support.</p>
<p>Pain is where it starts. Something hurts — revenue, reliability, speed, customer experience — and that pain is the signal that a real problem exists. But pain alone isn’t a goal. In product engineering, we shape pain into a problem statement, and the problem statement is what drives the work. The thematic goal has to be built on a problem that’s genuinely shared — one that every team can feel in their own work, not just observe in someone else’s. If the goal doesn’t map to a problem that teams are already living with, it won’t generate the shared investment you need. It’ll generate compliance at best.</p>
<p>Across the ten teams in our area, “break the monolith” was a problem every one of them felt directly — slow deployments, risky changes, coordination overhead. The goal named the problem. That’s what made it land.</p>
<h2>2. Active Coalition-Building</h2>
<p>Setting the goal is not leading the goal.</p>
<p>This is the part that distinguishes a thematic quarter from a well-named OKR exercise. The goal needs to be the question in every relevant forum. VPs and directors need to be actively mapping which teams need to talk to each other and making those introductions rather than assuming coordination will happen organically. Cross-team touchpoints need to be <em>created</em> — not just a Slack channel, but actual recurring forums where teams report blockers to each other, not just upward.</p>
<p>My <a href="https://blog.dicko.dev/posts/the-org-chart-trap-why-your-company-structure-is-silently-killing-velocity/">Org Chart Trap post</a> argues that Conway’s Law creates an almost gravitational pull toward the shape of your reporting lines. A thematic goal is one of the few tools that can temporarily override that gravity. But it can’t do it by being announced. It has to be the thing the people with organisational authority are actively using that authority to push.</p>
<p>My <a href="https://blog.dicko.dev/posts/breaking-engineering-silos-the-uncomfortable-truth-about-collaboration/">silos post</a> covers why silos form. This is how you build the temporary bridge across them. The bridge has to be maintained, repeatedly, by the people with the authority to hold it open.</p>
<h2>Running a Thematic Quarter: The Practical Version</h2>
<p>If you want to try this, here’s what I’d suggest based on what’s worked and what hasn’t.</p>
<p><strong>Set one goal. One.</strong> Not “our themes for the quarter.” One thing the whole leadership layer is trying to accomplish. It should be specific enough that a team could argue, convincingly, that what they’re working on either contributes to it or doesn’t. Vague goals survive the announcement meeting and die everywhere else.</p>
<p><strong>Make sure it hurts everyone.</strong> Before you announce it, stress-test it. If you’re small enough, talk to every team lead directly and ask what they’d have to deprioritise. If you’re larger, sample deliberately — a few team leads from diverse areas, not just the ones closest to you. At scale, a short survey works well: a single question asking what teams would have to stop or slow down if this became the priority. If the answer coming back is “nothing we weren’t already doing,” the goal isn’t thematic. Keep sharpening until everyone has to give something up.</p>
<p><strong>Appoint a coalition builder.</strong> Someone whose job — for the duration of the quarter — is to actively map and facilitate the cross-team conversations the goal requires. This isn’t a passive role. It’s making introductions, attending cross-team reviews, and noticing when two teams are solving the same problem from different angles without knowing it.</p>
<p><a href="https://blog.dicko.dev/posts/the-signature-question-how-culture-actually-spreads/"><strong>Ask the question obsessively.</strong></a> In every planning session. In every 1:1 with team leads. In every technical OKR review. <em>“How is this helping with [the goal]? If it’s not, why are we doing it?”</em> Not as a gotcha. As a genuine navigation question. If the answer is never good enough to stop asking, it means the goal is working. If the question starts feeling rhetorical, the goal is drifting.</p>
<p><strong>Create a cross-team visibility forum.</strong> At ten teams, a monthly demo works. Past fifteen, you probably need to structure this differently. But some mechanism where teams can see each other’s progress, surface blockers to each other rather than just upward, and feel the shared accumulation of work — that mechanism is load-bearing.</p>
<p><strong>Measure the right thing.</strong> The OKR for the thematic goal should measure the outcome you’re trying to achieve, not the effort you’re putting in. And make sure the measurement makes the work feel meaningful: the monolith migration tracked MRs moving to new services and celebrated each one. Progress should be visible to the people doing the work, not just the people reviewing it in a quarterly business review.</p>
<h2>The Bottom Line</h2>
<p>Most organisations set OKRs as a goal-setting exercise. The best ones use them as a coordination mechanism. The difference is whether there’s a single shared objective with enough gravity to temporarily override the gravitational pull of team boundaries — and whether someone with authority is actively building the coalition to make that override real.</p>
<p>Lencioni’s insight, filtered through a few years of actually trying it: the thematic goal is a manufactured crisis. It works because it creates the same shared context that a real crisis creates — everyone knows what matters, everyone knows why, and protecting territory stops feeling worth it. But unlike a real crisis, you can choose when it starts, what it’s about, and how much it costs.</p>
<p>The hard part isn’t the framework. The hard part is setting a goal that actually hurts everyone enough to matter, and then not abandoning it when the quarter gets busy and it becomes easier to let teams drift back to their local priorities.</p>
<p>The even harder part is asking the same question, in every room, for twelve consecutive weeks. By week four it feels redundant. By week eight it starts to feel like leadership. By week twelve you’ll have learned whether the goal had legs — or whether you announced it and hoped.</p>
<p>Now, if you’ll excuse me, I have a 1:1 in fifteen minutes and I need to think of a diplomatic way to ask “how is this helping with the goal” for the eleventh consecutive week. Apparently I’m still working on making it sound fresh.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Your New System Is Done. Now Good Luck Getting 40 Teams to Care.</title>
    <link>https://blog.dicko.dev/posts/your-new-system-is-done-now-good-luck-getting-40-teams-to-care/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/your-new-system-is-done-now-good-luck-getting-40-teams-to-care/</guid>
    <pubDate>Sat, 14 Mar 2026 05:21:34 GMT</pubDate>
    <category>engineering-leadership</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>api-design</category>
    <category>microservices</category>
    <description>Or: What Nobody Tells You About Retiring a Legacy System</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/your-new-system-is-done-now-good-luck-getting-40-teams-to-care/cover.webp" alt=""></figure>
<p>Somchai had been quiet for about three seconds — long enough that someone had started talking about the next agenda item — when he looked up at the whiteboard and asked the question. The kind of pause that happens when someone is choosing their words carefully, not because they don’t know what to say, but because they know exactly what they’re about to do to the conversation. The whiteboard still had the migration plan on it in confident blue marker — arrows from the old system to the new one, client boxes lined up like dominoes ready to fall. The meeting had been going well. Everyone had agreed the new system was good. The plan made sense. And then Somchai said: <em>“Have you considered just implementing the old system’s contracts in the new one so clients can change a URL and nothing else breaks?”</em></p>
<p>The whiteboard didn’t change. But the arrows on it suddenly looked different.</p>
<p>Let’s talk about the plan that nobody argues with in the planning meeting, but that quietly costs you years of your life anyway.</p>
<p>Here’s the shape of it. You build a new system. It’s good — maybe the best thing your team has shipped in years. Clean contracts, sensible design, the architecture you’d have built the first time if you’d known then what you know now. The plan: migrate all internal clients across. How hard can it be? You send out the announcement. You write the migration guide. You schedule the deprecation date.</p>
<p>And then you wait.</p>
<p>Some teams move fast. A few move in the first month, enthusiastic early adopters who actually read the announcement. Then the pace slows. Client F has a frozen deployment window until Q3. Client G’s team is two people down and won’t be able to prioritise this until after the rebrand. Client H’s lead just left, nobody’s quite sure who owns it now, and the deprecation notice is sitting unread in a team inbox that may or may not be actively monitored.</p>
<p>A year passes. You’re still supporting both systems.</p>
<p>I watched a friend live through this. A new API — well-designed, the right thing, clearly superior to what it replaced. The plan was standard playbook: build new thing, migrate clients, retire old thing. It took six years.</p>
<p>Six years of keeping the old system alive. Six years of patches, of incident reviews, of onboarding new engineers into the lore of a system that was <em>supposed to be dead</em>. By year four, my friend was having the same conversation on rotation — a gentle reminder email here, an escalation there, a meeting where someone who’d been cc’d on the original announcement said they weren’t sure this had ever been communicated to their team. In one of those meetings, I told him — cheekily, in the way I’d been saying it for years to nudge teams toward inner sourcing — “I accept pull requests.” The implication being: your migration, your pull request. His response: his team didn’t have the capacity to do the migration work across fifty client systems.</p>
<p>Let that sit for a moment. The team that created migration work across fifty client systems didn’t have capacity to do it themselves — so the plan was that fifty other teams would each find the time to do it for them. Nobody was the villain. Everyone was busy with their own priorities. <strong>That’s exactly the problem.</strong> You can’t build a migration plan whose success depends on fifty teams making your work their priority. That’s not a plan. It’s a wish.</p>
<h2>The Maths Nobody Does at the Start</h2>
<p>The instinct that drove the six-year migration is almost universal. The moment the new system is built, the team is mentally done. The old system is legacy — a liability. The plan is to kill it as fast as possible. So the default playbook is: announce, document, set a deprecation date, hope.</p>
<p>What almost nobody does is price the alternative up front.</p>
<p>The arithmetic isn’t complicated:</p>
<p><strong>Option A — Hard cutover, clients own their migration:</strong></p>
<ul>
<li><p>Build the new system- Keep the old system alive while clients migrate — weeks, months, or years, depending on how many clients you have and how much they care- The ongoing support cost of the old system runs for exactly as long as the slowest client takes to move
<strong>Option B — Build a compatibility surface in the new system:</strong></p>
</li>
<li><p>Build the new system plus a layer that speaks the old system’s contracts- The old system can be decommissioned as soon as the new one has feature parity and compatibility — you’re no longer blocked on client timelines- Clients migrate quickly — it’s just a URL change; the question is whether implementing the old system’s contracts in the new one costs less than supporting the legacy system for however long that tail turns out to be without it.
The crossover question is simple: is the cost of building the compatibility layer less than the cost of supporting the old system for the duration of the tail? If you’re migrating three internal clients you fully control, Option A is almost always right. If you’re migrating fifty clients with independent deployment windows, external vendors, and teams who will simply deprioritise your migration quarter after quarter — the maths turns against you fast.</p>
</li>
</ul>
<p>The variable nobody prices honestly is migration fatigue. Chasing clients to update a config is not free. Every follow-up email, every escalation, every meeting where someone has to explain again why this is important — that cost is real and it lands on the team that’s doing the retiring. It doesn’t show up in the project estimate. It shows up in engineer morale, three years later, when the person responsible for the migration has moved on and the old system is still running because it’s load-bearing in ways nobody fully mapped.</p>
<h2>The Pattern Has a Name</h2>
<p>Eric Evans named it in <em>Domain-Driven Design</em> in 2003: the Anti-Corruption Layer. His original framing was defensive — when a new system must integrate with a legacy system whose model is messy, you build a translation layer to stop the mess leaking into your clean new design. It’s the architectural equivalent of a quarantine zone.</p>
<p>The migration variant inverts the direction. Instead of the new system protecting itself <em>from</em> the legacy, the new system <em>mimics</em> the legacy’s contracts so that its clients don’t have to change. The ACL faces outward, speaking the old language. The internals stay clean.</p>
<p>When Somchai raised this in our meeting, he didn’t call it an anti-corruption layer. He said: <em>“implement the deprecated system’s contracts in the new system so clients can just change the URL.”</em> The vocabulary came later when someone asked what this pattern was called. The plain English was what actually landed.</p>
<p>This is worth naming because there are people in your organisation right now who would recognise the description immediately and have never once connected it to the DDD term. The concept is intuitive once you hear it. The vocabulary is what makes it reach for-able — the difference between a tool you can grab when you need it and an insight you can only articulate after the fact.</p>
<p>It’s also worth distinguishing from the Strangler Fig pattern, because they’re constantly conflated. Martin Fowler’s Strangler Fig is about gradual replacement through routing — you proxy traffic from old to new, migrating feature by feature, until the old system is strangled out of existence. The ACL compatibility variant is different: the new system is already built. The question is whether its external contracts match what clients already speak. One is a routing mechanism. The other is a compatibility surface. They’re complementary, not synonymous.</p>
<h2>Why the Room Goes Quiet</h2>
<p>Here’s the thing about the moment Somchai asked his question. He owned one of the client systems. He was, effectively, one of the dominoes on that whiteboard.</p>
<p>He wasn’t an outside architect spotting a missed option in someone else’s plan. He was the person about to bear the migration cost, suggesting a smarter way to absorb it. What he was really asking was: <em>before I go change my system, have you considered building a compatibility surface so I don’t have to?</em></p>
<p>The room went quiet because it forced everyone to ask a question the plan hadn’t asked: <em>how long is this migration actually going to take, and who’s paying for the support cost while it happens?</em></p>
<p>The economics professor Thomas Sowell has a line that applies here with uncomfortable directness: <em>“There are no solutions. There are only trade-offs.”</em> The hard cutover isn’t free — it just hides the cost in the migration tail. The ACL isn’t free either — it hides the cost in the upfront build. The question is which cost you’d rather pay, and when. The plan that doesn’t ask the question is the plan that defaults to the hidden cost every time.</p>
<p>In this particular case, the ACL wasn’t built. Three clients — small enough that the hard cutover was the right call. The maths worked out the other way, which is exactly the point. The ACL isn’t always the answer. The point is to do the maths rather than default to the standard playbook without asking.</p>
<h2>The Second-Order Risk Nobody Plans For</h2>
<p>There’s a coda to this that I find darkly funny, because it happened.</p>
<p>One ACL implementation at Agoda was named — informally, in the code — the “shitification layer” by the engineers who built it. The name is accurate: it made the new system speak the old system’s language, and the old system’s language was not something anyone was proud of. The pattern worked. The old system was retired. Mission accomplished.</p>
<p>The shitification layer is still there.</p>
<p>The ACL is not a free lunch. If you build one, it is not self-cleaning. It needs an owner. It needs a sunset date that someone is accountable for. It needs to be treated as first-class infrastructure — not a convenient bridge you build and then forget about. The trap is that once you’ve absorbed the migration cost into the new system, the urgency to actually complete the migration drops to nearly zero. Clients have no pressure to move. The old contracts keep working. Months become years. Your clever architectural move becomes the next legacy system.</p>
<p>The surgeon and writer Atul Gawande wrote in <em>The Checklist Manifesto</em> that the critical failures in complex systems almost never happen because someone didn’t know what to do — they happen because the right question wasn’t asked at the right time. “Have you considered building an ACL?” is one of those questions. “What’s the exit strategy for the ACL?” is the other one. Ask both, up front, or you’ve only solved half the problem.</p>
<h2>What Changes If You’ve Read the Book</h2>
<p>There’s a quieter point underneath all of this, and it’s the one I keep coming back to.</p>
<p>Every person in that meeting with Somchai was experienced. Smart. Good at their jobs. None of them had reached for the ACL option before he raised it — including me. I’d lived the six-year migration. I knew viscerally why the standard playbook fails. But I hadn’t connected the experience to the vocabulary, which meant the tool wasn’t available when I needed it.</p>
<p>Somchai had read <em>Domain-Driven Design</em>. Not as a box to check — actually read it. And when he needed the frame, it was there.</p>
<p>This is not an argument that reading is better than experience. It’s an observation that they’re different paths to the same insight, and only one of them is transferable. The six-year migration made me a better engineer. It stayed with me. It will inform every system retirement conversation I have for the rest of my career. But it didn’t transfer — it didn’t help anyone in that room until I was able to tell the story. The vocabulary transfers. That’s the reason I write this post.</p>
<h2>What to Actually Do</h2>
<p>If you’re building a new system that’s replacing something with multiple clients:</p>
<p><strong>Before the design is finalised</strong>, ask: how many clients does the old system have, what is their realistic migration timeline, and what is the weekly cost of supporting the old system while they migrate? Do the maths. Write it down. Make the trade-off explicit.</p>
<p>The equation is simple:</p>
<blockquote>
<p><em><strong>Cost of implementing the old contracts in the new system</strong></em></p>
</blockquote>
<blockquote>
<p><em>vs.</em></p>
</blockquote>
<blockquote>
<p><em><strong>Weekly cost of supporting the old system × number of weeks until the last client migrates</strong></em></p>
</blockquote>
<p>If the left side is smaller, build it. If the right side is smaller, hard cutover. The mistake isn’t choosing wrong — it’s never doing the calculation at all.</p>
<p><strong>If an ACL is the right answer</strong>, treat it as first-class infrastructure from day one. Give it an owner — not the “person closest to it” but an actual named owner with accountability. Set a sunset date with teeth: a date after which the compatibility contracts will actually be removed, communicated far enough in advance that clients have time to move. If the sunset date passes without action, you haven’t built an ACL — you’ve built a second legacy system.</p>
<p><strong>When you’re in the meeting</strong> and someone asks the question that makes the room go quiet — don’t reach for the whiteboard immediately. Sit with it. The discomfort you’re feeling is the plan being stress-tested, and that’s exactly what it needs.</p>
<h2>The Bottom Line</h2>
<p>Every system replacement project feels like it’s mostly a technical problem with a small coordination overhead. The ones that take six years felt that way too.</p>
<p>The migration tail is where plans go to die slowly. It’s not dramatic. Nobody calls a post-mortem on a migration that took four years longer than estimated — there’s no incident, no outage, just a quiet accumulation of ongoing cost that never made it onto the original estimate. The old system is still running. The engineer who knows how it works just accepted a job somewhere else. The deprecation announcement is three years old and nobody reads three-year-old announcements.</p>
<p>You don’t have to earn this lesson the hard way. The maths isn’t complicated. The pattern exists. The question — <em>“what if the new system just spoke the old contracts?”</em> — is available to anyone willing to ask it.</p>
<p>Ask it before you start.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Knowing It’s Fixed: How to Actually Validate a Performance Improvement</title>
    <link>https://blog.dicko.dev/posts/knowing-its-fixed-how-to-actually-validate-a-performance-improvement/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/knowing-its-fixed-how-to-actually-validate-a-performance-improvement/</guid>
    <pubDate>Mon, 09 Mar 2026 09:55:13 GMT</pubDate>
    <category>web-performance</category>
    <category>frontend-development</category>
    <category>ab-testing</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <description>Or: Shipping a Fix Is Not the Same as Knowing It Worked</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/knowing-its-fixed-how-to-actually-validate-a-performance-improvement/cover.webp" alt=""></figure>
<p>Namfon had shipped the new page on a Tuesday. The vinyl floor on level 6 was catching the afternoon light the way it does when the day has gone well — that particular low-angle Bangkok gold that comes through the floor-to-ceiling glass around 4pm — and the dashboard looked exactly right. LCP was green. Load time was green. She’d built the thing, measured the thing, and the numbers confirmed what she already suspected: it was fast. She marked it done and went home.</p>
<p>The complaints started the following week. Not a flood — a trickle, the kind that’s easy to dismiss as users being users. The page felt slow. Something wasn’t loading. It took ages to see anything useful. We pulled up the dashboards. Everything was green. We pulled up the RUM data. Still green. For a day or two, we assumed the users were wrong.</p>
<p>Then someone opened Chrome’s performance tooling and actually watched the LCP marker fire. It landed at 800 milliseconds — crisp, fast, exactly what you’d want. The problem was what was on screen at 800 milliseconds: the header. The navigation bar. A page shell that looked like a page but contained nothing a user could act on. The data table — the entire reason anyone opened this page — arrived 1,300 milliseconds later. On the office network. For real users on real connections, it was worse. We’d been measuring the arrival of the furniture and calling it the arrival of the meal. The compass had said we were on course. We weren’t anywhere near the destination.</p>
<p>This is the third and final post in this series. <a href="https://blog.dicko.dev/posts/your-performance-dashboard-is-lying-to-you/">Post 1 covered instrumentation — how to measure the right things for your actual users, not just the metrics your tooling defaults to.</a> <a href="https://blog.dicko.dev/posts/where-web-performance-actually-lives/">Post 2 covered intervention — where to look for the real performance levers in a modern frontend stack.</a> This post covers the part nobody talks about honestly: how do you know the fix worked?</p>
<h2>Why CI Can’t Close the Loop</h2>
<p>The obvious answer is CI performance benchmarks — fail the build if LCP exceeds a threshold, same as you’d fail it for test coverage. And CI benchmarks are worth having. They catch egregious regressions before they reach production. They’re a useful signal.</p>
<p>But they have a fundamental limitation that’s worth being honest about: <strong>CI cannot replicate production conditions.</strong></p>
<p>The factors that actually determine real-user performance — device diversity, network variance, CDN cache state, concurrent server load — are all invisible to CI. Your benchmark runs on one machine, with one network, with a clean cache.</p>
<p>If you’re running A/B tests, the problem compounds. Your CI pipeline tests variant A — the control, the safe default. The change you’re trying to validate is behind a feature flag that’s off by default. Your Lighthouse score looks pristine. Your synthetic test is green. Production is not your CI environment.</p>
<p>There’s a practical workaround: force a single variant in CI and run your performance suite twice on high-priority changes — once with all-A variants, once with the PR’s variant as B. Compare the two runs. This doubles CI time for those runs, so you do it selectively, not universally. It’s a useful regression detector. It is not a reliable answer to “did this help?”</p>
<h2>A/B Testing Is the Validation Layer</h2>
<p>Here’s the thing most teams don’t realise: if you’re already running A/B tests for product experiments, you already have the infrastructure to validate performance work. You just aren’t using it that way.</p>
<p>The only reliable method for knowing whether a performance change improved the real-user experience is to show the change to a slice of real users and compare their experience to a control group simultaneously, under the same conditions, at the same time. That is exactly what A/B testing does.</p>
<p>There’s a subtlety here worth stating explicitly: if you’re running multiple experiments concurrently — and at any mature product company, you are — a post-deployment snapshot can’t isolate your change from everything else happening in production. An A/B test can. Because both variants see the same traffic, the same device mix, and the same concurrent experiments at the same 50/50 split, any external factor affects A and B equally. It cancels out. What’s left is the signal from your change alone. That’s the property that makes it reliable in a way no post-deploy chart can match.</p>
<p>The leap is treating performance metrics the same way you treat conversion or engagement: as tracked outcomes in the experiment. If your experimentation platform can split users across variants — and it can — it can split them across “old code” and “performance-improved code.” Add LCP, <a href="https://web.dev/articles/inp">INP</a>, or your custom metrics as outcomes alongside your product metrics. Now you have a simultaneous comparison with real users under real conditions, not a synthetic test on a single machine.</p>
<p>At Agoda, we integrated Web Vitals telemetry directly into our A/B testing system. Every experiment running in the product produces a per-variant performance comparison automatically. The effect of this is twofold: performance regressions in new features are caught at experiment ramp-up, before full launch. And performance improvements ship with evidence — actual numbers from real users, not a favourable post-deployment snapshot that might be explained by a traffic shift you didn’t account for.</p>
<p>The honest cost is time. A/B tests need to reach statistical significance before you can draw conclusions. A small improvement on a low-traffic page might take weeks to produce a confident result. That is the trade-off: you exchange speed of certainty for accuracy of certainty. For a team that’s been making performance decisions based on charts that might be telling them what they want to hear, it’s a trade worth making.</p>
<h2>The Regression You Won’t See</h2>
<p>Standard Web Vitals — LCP, INP, CLS — are well-validated metrics for consumer web contexts. For B2B or data-heavy applications, a meaningful regression can be completely invisible to all three, while significantly degrading the experience of the users the product exists to serve.</p>
<p>The scenario: you ship a change that makes a backend data query two seconds slower. The page has a header, a navigation bar, and a large data table.</p>
<ul>
<li>LCP fires when the header — the largest element in the initial render — has painted. The data hasn’t arrived yet. LCP reports clean.- INP is unaffected. Click responsiveness hasn’t changed.- CLS is unaffected. Nothing shifted.
Every standard metric is green. The user is waiting two extra seconds for the data they opened the page to act on. The regression exists; it’s invisible to the instruments you’re using. See real below example:</li>
</ul>
<figure><img src="https://blog.dicko.dev/posts/knowing-its-fixed-how-to-actually-validate-a-performance-improvement/image_2.webp" alt="" loading="lazy"></figure>
<p>This is the problem that motivated what we call the Critical Mark — a developer-nominated readiness marker that fires when the primary data component mounts. Not when the page shell loads. When the content the user needs is actually present.</p>
<p>On one of our extranet product pages, LCP fires when only the page header has loaded. The data table — the content the page exists to provide — isn’t there yet. LCP reports a healthy metric. Critical Mark, attached to the data component, fires significantly later. The gap between those two numbers is the measurement error that would have been invisible without the custom instrumentation.</p>
<p>This is the case for custom performance metrics: not as a replacement for standard ones, but as a supplement that covers the gap between “page rendered” and “user can do their job.” As we covered in <a href="https://blog.dicko.dev/posts/your-performance-dashboard-is-lying-to-you/">Post 1 of this series</a>, the right metric is the one that reflects actual user workflows in your actual application — and sometimes you have to build it yourself.</p>
<h2>Disciplined Restraint: Knowing When Not to Ship</h2>
<p>The most mature version of “knowing it’s fixed” is recognising when you don’t have enough evidence to act.</p>
<p>We ran proof-of-concept work on <a href="https://developer.chrome.com/blog/search-compression-dictionaries">Compression Dictionary Transport</a> (CDT) — a newer HTTP standard that uses a browser’s cached version of your previous bundle as a compression dictionary to dramatically reduce the payload size of subsequent deployments. The technology is real and validated at scale: <a href="https://developer.chrome.com/blog/search-compression-dictionaries">Google Search reported around 23% payload reduction</a> when adopting it. The promise for teams with frequent deployments is significant — each deploy is cheaper for users who already have an older bundle cached.</p>
<p>The expected signal for CDT benefit is a performance dip around deployment time: the window when users first encounter a new bundle version before their browser has built the dictionary from the old one. We looked for this signal in our aggregated performance data. We couldn’t clearly see it. The dip either exists but isn’t visible through our current data aggregation approach — we’d need to look at a much tighter time window around specific deployments rather than daily aggregates — or Agoda’s specific deployment patterns don’t produce the problem CDT would solve at the scale the technology’s benefits would justify.</p>
<p>The team’s position: find the problem clearly before deploying the solution.</p>
<p>This is an underappreciated form of engineering discipline. The instinct, when promising technology appears, is to adopt it. CDT is good technology. The evidence that it works is credible. But “this technology works” and “this technology solves our specific problem” are different claims, and we hadn’t established the second one. Optimising into the unknown — shipping a performance improvement you can’t measure the effect of, hoping the numbers improve — is the validation failure in reverse.</p>
<p>As the statistician W. Edwards Deming put it: <em>“Without data, you’re just another person with an opinion.”</em> That applies as much to the decision to ship a performance optimisation as it does to the decision not to.</p>
<h2>Closing the Loop</h2>
<p>The three posts in this series describe a single loop:</p>
<p><strong>Instrument correctly.</strong> Real User Monitoring, the right metrics for your users and application type, at the right percentile for your noise floor. Custom markers where standard metrics don’t reach.</p>
<p><strong>Find the right lever.</strong> BFF parallelism, loading state sequencing, bundle composition. Fix the things that actually matter for your users, not the things that are visible in synthetic tests.</p>
<p><strong>Validate reliably.</strong> A/B testing as the production validation layer. Custom metrics as regression detectors. Disciplined restraint when you can’t see the problem clearly enough to act on it.</p>
<p>The loop doesn’t close unless all three work. The research on high-performing engineering teams — <a href="https://itrevolution.com/product/accelerate/">Forsgren, Humble and Kim’s <em>Accelerate</em></a> being the most rigorous — consistently shows that measurement practices are what separate teams that improve from teams that just stay busy. Teams that measure well but fix the wrong things waste engineering effort. Teams that fix the right things but can’t validate whether they worked are navigating with a broken compass. Teams that validate well but don’t have the right instrumentation will miss regressions that happen in the gap between “page loaded” and “user can work.”</p>
<p>Performance is not a feature you ship. It’s a practice you maintain — and like any practice, it only improves when you’re honest about what you actually know versus what you want to believe.</p>
<p>Back on level 6, the dashboards were still green when we found the problem. Namfon hadn’t done anything wrong — she’d measured exactly what the tooling told her to measure, and the tooling had been measuring the wrong thing. What the Chrome performance trace gave us wasn’t a fix — it was the discovery that our instrumentation had been lying to us, probably on more pages than just this one, and that fixing it would be worth more than any individual performance win. We built the Critical Mark. We added performance outcomes to our experiment results. We stopped calling things done until we were sure we were measuring what users actually experienced.</p>
<p>The compass works now. Mostly.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Where Web Performance Actually Lives</title>
    <link>https://blog.dicko.dev/posts/where-web-performance-actually-lives/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/where-web-performance-actually-lives/</guid>
    <pubDate>Sun, 08 Mar 2026 13:46:32 GMT</pubDate>
    <category>web-performance</category>
    <category>frontend-development</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>core-web-vitals</category>
    <description>Or: Why Your Frontend Engineer Can’t Fix What Isn’t a Frontend Problem</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/where-web-performance-actually-lives/cover.webp" alt=""></figure>
<p>Namfon had done everything right. The Lighthouse score sat at 98. Images were lazy-loaded, the critical CSS was inlined, render-blocking scripts had been hunted down and eliminated. The lab numbers were clean. But the field data — real users, real devices, real networks — told a different story, and it had been telling it for weeks.</p>
<p>It was around 11am on a Wednesday, the kind of Bangkok morning where the AC on level 6 was already losing the argument with the sun coming through the floor-to-ceiling glass. The open-plan floor hummed — standups bleeding into each other from three directions, someone’s Teams notification going off every forty seconds, the faint plastic smell of fresh Thai tea from the kitchen drifting past. Namfon’s desk had that specific archaeology of someone who hadn’t quite left a problem: a cold cup of something she’d forgotten to drink, three browser tabs frozen mid-audit, a sticky note with “WHY” written on it and nothing else. She leaned back from her monitors, crossed her arms, and said to no one in particular, “I don’t know where else to look.”</p>
<p>She wasn’t wrong. She just wasn’t looking in the right place — because the right place wasn’t in the browser at all.</p>
<p>This is the second post in a three-part series on web performance. <a href="https://blog.dicko.dev/posts/your-performance-dashboard-is-lying-to-you/">Post 1</a> covered how to <em>see</em> the problem — the metrics, the instrumentation, the specific incident that made me take INP seriously. This post is about where the fix actually lives once you can see it.</p>
<p>The short answer: usually not in the browser.</p>
<p>The browser is the surface where milliseconds are <em>felt</em>. It’s where Lighthouse runs, where your product manager takes screenshots of loading spinners, where the user’s frustration lives. But in a modern microservices architecture, the browser is often just the messenger — faithfully reporting the outcome of decisions made much further upstream. Decisions about how your BFF aggregates data, how your bundles are structured, and whether your teams own performance or just blame each other for it.</p>
<p>Let’s go through where the leverage actually is.</p>
<h2>The BFF Layer: Sequential vs Parallel Is the Highest-Leverage Intervention You’re Probably Not Making</h2>
<p>A rich page in a microservices architecture needs data from multiple services. User preferences, pricing, availability, promotional content — each comes from a different place. The Backend for Frontend (BFF) pattern exists to aggregate these into a single, page-optimised response, so the browser makes one clean API call rather than juggling five or ten.</p>
<p>But here’s the thing about that BFF call: the way it’s written internally determines whether your page loads in 100ms or 500ms, and it usually has nothing to do with any of the individual services.</p>
<p>If the BFF fetches those services sequentially — <code>await serviceA()</code>, then <code>await serviceB()</code>, then <code>await serviceC()</code> — the page response time is the <em>sum</em> of all service latencies. Five services at 100ms each: 500ms. If it fetches them in parallel, the response time is the <em>maximum</em> of all service latencies. Five services at 100ms each: 100ms.</p>
<p>Same services. Same backend. Same frontend. Five times faster.</p>
<p>The fix is small code: <code>Promise.all([serviceA(), serviceB(), serviceC()])</code> in JavaScript, <code>Task.WhenAll()</code> in C#. The reason it doesn&#39;t get written this way the first time is simple — sequential awaits are easier to read and reason about. You write what makes sense in the moment, and the performance implication is invisible until you have real load and real data.</p>
<p>The real BFF, of course, is more nuanced than “just parallelise everything.” Some calls have dependencies: service B might genuinely need the result of service A before it can run. The right mental model isn’t “full sequential” or “full parallel” — it’s a dependency graph. What can run independently of what? Group the independent calls together, chain only where you actually have to. That analysis — mapping your BFF call graph and identifying what’s blocking unnecessarily — is one of the most underappreciated architectural tasks in frontend engineering, and it’s almost always done either too late or not at all.</p>
<p>One more thing that often gets missed before we even get to the BFF: browsers enforce a hard limit of six parallel HTTP/1.1 connections per domain. If your API calls and your static assets are competing for the same six slots, they’re throttling each other at the browser level. Serving your static assets — JS bundles, CSS, images — from a dedicated CDN domain frees those connection slots entirely for API traffic. This looks like a deployment detail. It’s actually a browser constraint optimisation, and it’s the reason you want to solve this at the BFF before you even think about making the frontend juggle multiple service calls directly.</p>
<p>Our BFFs at Agoda are .NET (C#) and Kotlin — not for any philosophical reason, but because both runtimes handle concurrent aggregation work at scale more efficiently than Node’s single-threaded event loop. When your BFF is doing significant parallel I/O and aggregation simultaneously, the runtime choice matters. Sam Newman’s original <a href="https://samnewman.io/patterns/architectural/bff/">BFF pattern write-up</a> calls out parallel orchestration as a core concern of the pattern, not an optimisation to layer on later. We’ve found that’s correct.</p>
<h2>The Loading State You’re Getting Wrong (And Why It Matters at Scale)</h2>
<p>This one is small, specific, and embarrassingly common.</p>
<p>INP — Interaction to Next Paint — measures the time between a user doing something and the browser visually acknowledging it. The key word is “visually.” INP doesn’t measure how long your async work takes. It measures how quickly the <em>next paint</em> after interaction happens.</p>
<p>The correct sequence when a user clicks a button:</p>
<ul>
<li><p>User clicks- Button immediately shows a loading state — <em>this is the next paint INP measures</em>- Async work begins
The common (wrong) sequence:</p>
</li>
<li><p>User clicks- Async work begins- Button shows loading state after the first state update resolves
If async work starts before the DOM update, INP measures the entire duration from click to visible response — which includes the async work itself. A loading spinner that appears 300ms after click because it’s set inside a <code>.then()</code> is 300ms INP, not near-instant feedback. The fix is a state update <em>before</em> the async call, not inside it.</p>
</li>
</ul>
<p>This matters beyond user experience scores. At significant traffic volumes, a page that doesn’t visibly acknowledge clicks creates re-click behaviour. Users interpret a non-responsive button as “my click didn’t register” and click again. Each re-click is another request. On expensive operations — search queries, booking submissions — a mediocre INP score can turn a slow-but-functional page into an overloaded-and-down page under the right load conditions. The web.dev <a href="https://web.dev/articles/inp">INP documentation</a> is the right reference here, and it’s worth actually reading rather than just targeting the number.</p>
<h2>The DOM Nobody’s Looking At</h2>
<p>Here’s one that comes up repeatedly in INP investigations and almost never gets caught in a Lighthouse run: engineers rendering DOM that users can’t see.</p>
<p>It sounds absurd when you say it plainly. Why would you render a thousand rows when the user can only see twenty? But it happens constantly, and it’s almost always unintentional. A search results page renders the full list. A feed renders every item. A data table renders all rows up front because the original dataset was small and nobody noticed when it grew. The page looks fine. Scrolling feels fine. And then, at some point, React needs to reconcile a state change — a filter selection, a price update, a hover interaction — and it dutifully re-renders every node in that enormous DOM. Including the eight hundred rows the user can’t see because they’re below the fold.</p>
<p>INP spikes. Engineers look at the interaction code and find nothing obviously wrong. The component is clean. The logic is simple. The problem isn’t the code — it’s the surface area the code is operating on.</p>
<p>DOM size is a multiplier on everything. The larger the DOM, the more expensive every rerender, every style recalculation, every layout pass. Browsers aren’t infinitely fast at traversing and updating nodes, and React’s reconciliation, efficient as it is, still has to do work proportional to what’s mounted. Mounting less is always faster than reconciling faster.</p>
<p>The solution that actually works at scale is virtualisation — only rendering what’s visible in the viewport and dynamically creating and destroying nodes as the user scrolls. <a href="https://tanstack.com/virtual/latest">TanStack Virtual</a> is the library worth knowing here: it handles the windowing logic without prescribing your rendering approach, works with any framework, and manages the complexity of variable-height items that makes naive virtualisation implementations fall apart.</p>
<p>The pattern is straightforward in principle: you maintain a full list in memory, but at any given moment only a small window of that list exists in the DOM. As the user scrolls, nodes at the top get destroyed, nodes at the bottom get created. The user never notices. React rerenders become fast again because there are thirty nodes to reconcile instead of three thousand.</p>
<p>It won’t fix every INP problem. But if you have a slow interaction on a page with large lists or tables and you haven’t looked at DOM size yet, look there first. You might find the culprit isn’t the interaction at all.</p>
<p>As the statistician W. Edwards Deming observed, “Every system is perfectly designed to get the results it gets.” A high-tempo A/B testing culture is perfectly designed to accumulate JavaScript bundle bloat. Not because anyone is being careless — but because of how the system works.</p>
<p>Every A/B test requires both variants to be present in the bundle. The server doesn’t know at build time which variant a given user will see, so both ship. One test: a small amount of extra code. Twenty concurrent tests: a meaningful amount of dead code delivered to every user on every page load, where “dead” means “code the current user will never execute.”</p>
<p>This is the structural cost of experimentation speed. It’s not a mistake. But it compounds. Experiments conclude, their code should be removed, and in practice — honestly, in most organisations — it doesn’t get cleaned up immediately. The cleanup ticket sits in the backlog. The next sprint has higher priorities. The experiment code becomes part of the background noise, and then becomes load-bearing architecture because everyone’s forgotten what it was.</p>
<p>Three approaches exist for managing this:</p>
<p><strong>Discipline:</strong> Remove experiment code when tests conclude. Assign ownership. Make it a deployment blocker. This works at small scale. It degrades at high experiment velocity because the process burden scales with the number of experiments, not the team size.</p>
<p><strong>Compression Dictionary Transport (CDT):</strong> A relatively new HTTP standard, supported in Chrome and Edge from version 130, that uses a previously-cached version of a resource as the compression dictionary for the new version. Instead of downloading the full updated bundle on each deployment, the browser receives a compressed diff against what it already has. This doesn’t reduce what’s <em>in</em> the bundle, but it dramatically reduces what needs to be transmitted on update. Worth noting: CDT is most effective when you have regular returning users — B2B applications, social platforms, tools where daily or weekly users are common. For ecommerce, where a significant portion of users arrive for the first time or after long gaps, the benefit is reduced. The <a href="https://developer.chrome.com/blog/shared-dictionary-compression">Chrome developers overview</a> is the right starting point.</p>
<p><strong>Edge tree shaking per user:</strong> Knowing a user’s A/B experiment assignments at request time — which an edge worker can do — you can strip the dead-variant code before serving the bundle. Each user receives only the code their session will execute. This is the most aggressive solution and the most complex to implement. <a href="https://zephyr-cloud.io/">Zephyr Cloud</a> implements this as a bundler plugin plus edge worker and is worth examining if bundle size is a genuine bottleneck for you.</p>
<p>The honest position: most teams are running on discipline, watching it slowly fail, and haven’t yet seriously evaluated the alternatives.</p>
<h2>The GraphQL Upstream Problem</h2>
<p>We’re conditioned to optimise response payload size. Smaller responses download faster — this is obvious, it’s correct, and it’s incomplete.</p>
<p>HTTP connections are asymmetric. Significantly more downstream bandwidth is available than upstream. On mobile, particularly on inconsistent networks, upstream is a real constraint. In GraphQL-heavy architectures, the query body going <em>up</em> can itself be several kilobytes of JSON per request. On desktop, on a good connection, this is invisible. On mobile, at scale, it isn’t.</p>
<p>Our mobile teams at Agoda address this with request body compression on GraphQL queries — roughly 7:1 compression ratios on upstream payloads. Native mobile apps give you direct control over request construction, which makes this reasonably straightforward to implement. Web browsers don’t support request body compression the same way without explicit server negotiation, which makes this a mobile-first optimisation rather than a universal one.</p>
<p>The principle is worth holding onto regardless: in request/response cycles, both directions have cost. Your instinct to optimise response size is right. It’s just not the only thing worth optimising.</p>
<h2>Performance Is a Culture, Not a Team</h2>
<p>Here’s the failure mode that nobody wants to name because it feels unfair to the people working hard inside it.</p>
<p>Performance becomes a problem. Leadership creates a performance team to fix it. The performance team works hard, ships improvements, publishes dashboards. Meanwhile, every other engineering team continues making decisions — architectural choices, new features, third-party integrations, BFF call patterns — without performance as a first-class consideration. The performance team spends most of its time catching up with regressions introduced by everyone else.</p>
<p>Then the resentment sets in, from both sides. The performance team feels like they’re running up a down escalator. The other teams develop the attitude that kills this whole model: <em>“Don’t we have a team for that?”</em></p>
<p>As the surgeon and author Atul Gawande wrote about failure in complex systems: “The problem is rarely a lack of effort. The problem is that the system doesn’t make the right behaviour the easy behaviour.” Creating a dedicated performance team doesn’t make good performance decisions the easy behaviour for everyone else. It makes outsourcing those decisions the easy behaviour — which is the opposite of what you need.</p>
<p>Performance in a large engineering organisation is downstream of every decision every other team makes. Schema choices, component render patterns, BFF aggregation structure, experiment accumulation, API payload design — none of these are owned by a performance team. All of them affect performance. A team that <em>fixes</em> performance problems is a crutch. A team that <em>enables other teams to fix their own performance problems</em> is a multiplier.</p>
<p>The model that works looks like this: the performance team sets and maintains SLAs with product teams, so performance thresholds are commitments rather than aspirations. They build and maintain the tooling — RUM instrumentation, dashboards, custom metrics — so teams can see their own performance without needing the performance team’s involvement. They sit in on architecture reviews early, when the cost of changing a bad decision is low, not after the regression has shipped. They run training — not on Lighthouse scores, but on how browsers actually work, why the loading-state sequence matters, what sequential BFF calls are costing.</p>
<p>The goal is teams that know how to make their pages fast and keep them that way. Not faster pages that drift back to slow once the performance team moves on.</p>
<h2>What This Means in Practice</h2>
<p>Pulling back to the original scene: Namfon’s Lighthouse score was 98. The page was slow. Both things were true simultaneously, and understanding why requires holding two ideas at once.</p>
<p>The browser is where performance is experienced. It is not where performance is determined.</p>
<p>The BFF’s aggregation pattern determines whether you pay 100ms or 500ms before the browser has anything to work with. The bundle’s structure determines how much code a first-time visitor has to parse. The loading state sequence determines whether INP is measured in milliseconds or seconds. None of these show up as “frontend problems” on Lighthouse.</p>
<p>The fix Namfon needed wasn’t a better Lighthouse audit. It was a conversation with the team that owned the BFF — a look at the call graph, a question about what was running sequentially that didn’t need to be.</p>
<p>That conversation is harder than running a Lighthouse report. It crosses team boundaries. It requires shared ownership of a metric that nobody fully owns.</p>
<p>That’s precisely why it’s worth having.</p>
<p><em>Post 3 of this series covers how you know a fix is actually working — measurement, baselines, and the specific metrics that tell you whether you’re improving or just making a different kind of slow.</em></p>
]]></content:encoded>
  </item>
  <item>
    <title>The Autopsy Nobody Wants to Do</title>
    <link>https://blog.dicko.dev/posts/the-autopsy-nobody-wants-to-do/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/the-autopsy-nobody-wants-to-do/</guid>
    <pubDate>Sun, 08 Mar 2026 07:28:33 GMT</pubDate>
    <category>agile</category>
    <category>engineering-management</category>
    <category>software-engineering</category>
    <category>engineering-culture</category>
    <category>retrospectives</category>
    <description>Or: A Field Guide to the Execution Post-Mortem</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/the-autopsy-nobody-wants-to-do/cover.webp" alt=""></figure>
<p>Somchai was the first one in the room. It was 10:14 on a Thursday morning and the air conditioning on level 6 was doing its usual impersonation of effort — technically running, technically failing. He’d brought his laptop. He didn’t open it. He just sat at the end of the table and looked at the door, the way you look at a door when you know who’s coming through it and you know what they’re going to say. The room smelled of the Thai tea he hadn’t touched. The calendar invite had said “Milestone Review.” Everyone in the meeting knew that was a polite fiction. What it meant was: <em>the milestone was missed, and now we’re going to talk about it.</em></p>
<p>That meeting has a specific atmosphere. You’ve been in it. You know the particular quality of silence before it starts — not the comfortable silence of people thinking, but the loaded silence of people <em>preparing</em>. The product side rehearsing their version. The engineering side marshalling their context. Everyone in the room carrying a narrative they’ve been quietly sharpening since the milestone slipped.</p>
<p>What happens next is almost always the same. Someone says “failed.” Someone else says “it was complicated.” Someone mentions a dependency. Someone mentions the requirements changing. Someone’s voice tightens a little. Nobody’s lying, exactly — but nobody’s <em>right</em>, either. Because what’s happening isn’t a conversation about what happened. It’s a competition between memories.</p>
<p>This post is about how to stop that competition before it starts. Not by being a better mediator. Not by having better values or running better retros. But by walking into the room with something that memories can’t compete with: evidence.</p>
<h2>The Problem With the Conversation We Keep Having</h2>
<p>Most delivery post-mortems — when they happen at all — are conducted from memory. People describe their experience of the sprint. They talk about what felt slow, what felt blocked, what they were waiting on. And because everyone experienced the same period of time from a different vantage point, you get a beautiful collection of individually coherent, mutually incompatible accounts.</p>
<p>The loudest voice tends to win. The most senior person’s narrative tends to stick. The engineer who was quietly blocked for two weeks — blocked in a way that never made it to a standup, never landed in a ticket comment, never got escalated — stays quiet, because what are they going to say? “I was waiting”? Against someone with a slide deck?</p>
<p>Here’s the part that should genuinely unsettle you: in my experience running these investigations, the thing being blamed loudest in the room is almost <em>never</em> the actual constraint. The finger points at QA. The timeline shows QA was waiting two days for a build. The finger points at another team’s API. The timeline shows the dependency wasn’t tracked as a risk at planning and the team kept working as if it didn’t exist. The finger points at scope creep. The timeline shows the scope was defined precisely — the estimate was just wrong, and nobody re-evaluated it when reality diverged.</p>
<p>The culprits are almost always quieter. Internal silos. Analysis paralysis that looked like thoroughness. Large PRs opened after weeks of solo development that assumed alignment that hadn’t been established. Poor planning that the team knew was poor by day three but kept pushing through anyway, hoping.</p>
<p>You can’t see any of that in a retro. A retro is the wrong tool for this problem. A retro is two weeks of memory, fifteen post-it notes, and a timebox. A milestone failure usually spans <em>multiple sprints</em> and requires you to go back further than anyone’s accurate recall allows. What you need isn’t a retrospective. It’s a reconstruction.</p>
<p>That’s what I call the <strong>Execution Post-Mortem</strong>.</p>
<h2>What the Execution Post-Mortem Is</h2>
<p>An Execution Post-Mortem is a structured, evidence-based investigation into how a milestone actually unfolded — built from primary sources rather than memory.</p>
<p>Not feelings. Not narratives. Primary sources: ticket history, git logs, PR review timelines, deployment records, calendar data. Things that were written down at the time. Things that don’t shift depending on who’s telling the story.</p>
<p>The output is a <strong>Delivery Timeline</strong>: a chronological record of what happened to each piece of work, from the moment it entered the system to the moment it was done (or wasn’t). It shows you where time actually went. Day by day. Gap by gap.</p>
<p>Once you have the timeline, the conversation in that meeting room changes completely. Instead of “it was complicated,” you can say: <em>“Three stories sat in review for more than five days each — here’s what was happening during those gaps.”</em> Instead of “the requirements kept changing,” you can say: <em>“The story acceptance criteria changed after pointing, on this date, and we didn’t re-size or flag it — here’s what we’re doing differently.”</em></p>
<p>The timeline doesn’t assign blame. It assigns <em>ownership</em>. And ownership is something both product and engineering can actually work with.</p>
<p>As the journalist and author Katherine Boo once wrote about investigating complex human systems: <em>“The building of a narrative requires the destruction of comfortable assumptions.”</em> That’s exactly what a delivery timeline does. It destroys the comfortable assumption that you know what happened — and replaces it with what actually happened.</p>
<h2>Why It Works Psychologically</h2>
<p>Before we get to the steps, this is worth understanding: the method works not just analytically, but <em>socially</em>.</p>
<p>When you walk into that room with a timeline rather than a narrative, you change the dynamic. You’ve made the problem visible to everyone simultaneously, from the same data, at the same time. Nobody finds out what happened by listening to someone else’s account — everyone looks at the same picture together. That shift — from <em>testimony</em> to <em>evidence</em> — is what lets you have a collaborative conversation instead of a defensive one.</p>
<p>The young backend engineer who moved on to the next milestone while the frontend work sat with one person for three weeks? When the timeline surfaces that, his reaction isn’t defensiveness. It’s surprise. Because he genuinely didn’t know. He thought moving ahead was being efficient. Nobody had shown him otherwise. He was output-focused, not outcome-focused — and nobody had ever told him those two things could diverge. The timeline made the invisible visible, and the conversation that followed was coaching, not confrontation.</p>
<p>That’s the version of this meeting worth engineering toward.</p>
<p>But here’s the caveat, and it matters: the method is only as good as the conditions it’s run in. If the room isn’t psychologically safe, the same data that produces honest conversation in one place produces counter-accusations in another. The timeline is neutral. The people in the room are not. Your job, as the person running this, is to establish the conditions before the first data point goes on the board.</p>
<h2>Running It Blameless: What That Actually Means</h2>
<p>“Blameless” gets thrown around in engineering culture as if it’s self-explanatory. It isn’t. Here’s what it means in practice when you’re running an Execution Post-Mortem.</p>
<p><strong>No names on the board.</strong> When you’re building the timeline and presenting findings, you talk about <em>the story</em>, <em>the PR</em>, <em>the sprint</em> — never the engineer. Not “Somchai’s PR sat for five days” but “this PR sat in review for five days.” Not “the backend team didn’t flag the dependency” but “this dependency wasn’t flagged at planning.” The moment a name goes up, the person attached to it stops thinking about the problem and starts thinking about themselves. Everyone else in the room does the same.</p>
<p><strong>“We” owns the process. Nobody owns the mistake.</strong> The framing is always: what did our process allow to happen? A story sat unreviewed for five days because we don’t have a review SLA. A dependency was invisible because we don’t track them explicitly at planning. An engineer carried work alone for three weeks because we don’t have a practice of checking in on single-assignee stories. The team built the conditions. The team changes the conditions.</p>
<p><strong>Findings are systemic, not personal.</strong> The question is never “who made this decision?” It’s “what made this decision rational at the time?” Because almost every decision that looks wrong in hindsight looked reasonable when it was made. The engineer who built solo for three weeks wasn’t being reckless — he thought he was being efficient. The context that made that feel right is the thing worth examining.</p>
<p>None of this means individuals are never accountable. It means accountability lands on the right thing: what does this person need to do differently, and what does the system need to give them to make that possible? That’s a coaching conversation. It happens separately, privately, and after the post-mortem — not in the room.</p>
<h2>When the Room Goes Off-Script</h2>
<p>Even with all of that set up in advance, rooms don’t always cooperate. People are human, and delivery failures carry real frustration. Sometimes someone fixates on a person instead of a process and the energy in the room shifts from analysis to prosecution.</p>
<p>I’ve been in one of those. A very animated Italian engineer — loud, passionate, entirely convinced that the problem was a specific colleague who had made a mistake. He wasn’t wrong that a mistake had been made. He was completely wrong about why it mattered. He kept circling back to the individual, louder each time, while the actual process failure sat quietly in the data, unexamined.</p>
<p>I let him run for a bit. Then I said, very calmly: <em>“Don’t worry about him. After this meeting, we’re going to take him out the back and shoot him.”</em></p>
<p>He stopped mid-sentence. Then he started laughing — genuinely, loudly, the kind of laugh that releases something. When he came back up for air, he’d lost the thread of his rant. He looked slightly sheepish. And then we got back to the process.</p>
<p>I’m not recommending gallows humour as a universal facilitation technique. But the underlying move is worth understanding: you need a pattern interrupt, something that breaks the emotional momentum and creates a beat of distance. Humour works when the room already trusts you. A direct redirect works in other contexts — <em>“I hear you, and that’s a separate conversation. Right now I want to understand the system. Can we come back to the timeline?”</em> The specific tool is less important than the intent: you’re not dismissing the frustration, you’re redirecting the energy toward something that will actually produce a result.</p>
<p>The alternative — letting the room stay in blame mode — is how you end the meeting with one person feeling vindicated and everyone else feeling worse, and the process failure completely unaddressed.</p>
<h2>The Process: Step by Step</h2>
<h2>Step 1: Establish Your Anchors</h2>
<p>Start with two dates you can verify from a primary source — not from memory, not from “I think it started around”:</p>
<ul>
<li><strong>T₀</strong> — The date the first story for this milestone was added to a sprint. Not when the initiative was announced. Not when it was discussed in planning. When did it first land in an active sprint backlog?- <strong>T_end</strong> — Either the actual completion date, or the date it became objectively clear it wouldn’t complete on time.
Write these down. Every other data point you collect goes between these two anchors. You now have the edges of your picture.</li>
</ul>
<h2>Step 2: Map the PR Lifecycle for Every Story</h2>
<p>This is your primary data source. The git history, PR review history, and deployment record are all visible from the PR itself — and a PR has four dates that will give you 95% of the story:</p>
<ul>
<li><strong>First commit</strong> — when development actually started- <strong>Ready for review</strong> — when the engineer considered it done and opened it up- <strong>Approved</strong> — when the team signed off- <strong>Merged / in production</strong> — when it actually shipped
For a multi-sprint milestone you may be looking at a lot of PRs. Don’t try to investigate all of them at this stage. Just collect these four dates for each one, put them in a row, and calculate the gaps. That’s enough to see the shape of the problem. The gaps that stand out are the ones worth drilling into.</li>
</ul>
<p>What you’re looking for at a glance:</p>
<ul>
<li>A long gap between first commit and ready for review — was the engineer working in isolation for too long before seeking feedback?- A long gap between ready for review and approved — where was the bottleneck: waiting for a reviewer, review rounds, rework?- A long gap between approved and merged/production — change windows, deployment failures, manual steps that shouldn’t exist?
Most PRs will look fine. A few will have gaps that jump off the page. Those are the ones that need a second look. The four-date summary tells you <em>which</em> ones. Only then do you go deeper — into review comments, commit frequency, deployment logs — on the ones that earned it.</li>
</ul>
<p><strong>A note on spike stories:</strong> not everything in your milestone will have a PR. Spikes — investigation, design, and research stories — produce a document or a decision, not a commit. Don’t skip them. In fact, treat them with extra attention, because spikes are where analysis paralysis hides most comfortably. A spike that was estimated at two days and ran for two weeks won’t show up in any PR scan. It’ll just be a ticket that sat “In Progress” for a very long time while the team convinced themselves they were being thorough. For spikes, use the ticket history directly: when was it picked up, when was it closed, and what came out of it? If the answer to that last question is “another spike,” you’ve found something worth discussing.</p>
<h2>Step 3: Pull the Ticket History for Context</h2>
<p>The ticket history fills in what the PR can’t tell you: what happened <em>before</em> the first commit, and anything that happened outside the code.</p>
<p>For each story:</p>
<ul>
<li>When was it created in the backlog?- When was it added to the sprint?- When did it move to “In Progress”?- Any status reversals — stories moved back to “To Do” mid-sprint?- Any comments flagging blockers, dependency issues, or scope changes?- Any assignee changes?
The gap between “added to sprint” and “first commit” is the one to watch here. It’s often where planning failures live — a story that sat in a sprint for four days before anyone touched it, because it wasn’t actually ready to be worked on, or because the person who picked it up had three other things running.</li>
</ul>
<p>Most Jira instances keep a full audit trail. Pull it. The history was written at the time and doesn’t shift depending on who’s in the room.</p>
<h2>Step 4: Build the Table</h2>
<p>Take every data point and lay it out. For each story, you want the ticket anchor dates alongside the four PR lifecycle dates, with the gaps calculated between each. It should look something like this:</p>
<figure><img src="https://blog.dicko.dev/posts/the-autopsy-nobody-wants-to-do/image_2.webp" alt="" loading="lazy"></figure>
<p>If you have a whiteboard, draw it up as an actual timeline though, visualizing it like this for everyone to collaborate on works wonders, also using distance a time, its easy to compare, but make sure your scale is correctly proportionated to the time dimension, otherwise you’ll mask problems potentially.</p>
<p>UNF-114 and UNF-122 look fine. UNF-118 jumps off the page immediately. But so does UNF-116 — a spike that ran for sixteen days with no PR, no commit, nothing. The ticket just sat open. Whatever it was investigating, it took four times longer than it should have, and the stories that were presumably waiting on its outcome couldn’t start until it closed. That’s analysis paralysis made visible.</p>
<p>That’s your investigation target. Now you go deeper on those two — into the PR comments, the ticket history, the calendar — and leave the others alone.</p>
<p><strong>Total calendar days for UNF-118: 27. Original estimate: one sprint. Total calendar days for UNF-116: 16. Original estimate: three days.</strong></p>
<p>The question isn’t “why did these take so long?” in the abstract. For UNF-118 it’s three specific questions: why four days before the first commit? What happened in review for seven days? And what was blocking the merge for eleven days after approval? For UNF-116 it’s simpler and harder: what were we actually doing for sixteen days, and what decision came out of it?</p>
<p>That’s the investigation. The table doesn’t give you answers. It tells you <em>which questions to ask</em> — specific, answerable ones, grounded in data rather than retrospective impression.</p>
<h2>Step 5: Interrogate the Gaps</h2>
<p>For every gap that looks disproportionate relative to the story’s estimate, ask: <em>what was happening on those days?</em> A two-day gap on a one-point story is worth a question. The same two days on a thirteen-point story is probably noise. The threshold isn’t a fixed number — it’s your judgment about what looks out of proportion given what the work was supposed to be.</p>
<p>The timeline flags the gaps. The team explains them. Once you have the table in front of everyone, simply ask: <em>“UNF-116 was open for sixteen days against a three-day estimate — can someone walk us through what was happening there?”</em> You’ll get the answer in thirty seconds, and it will be more accurate than anything you’d reconstruct from Slack history or calendar data. The people who did the work know what happened. The timeline just gives them something specific to respond to, instead of a vague invitation to defend themselves.</p>
<p>Most gaps have an explanation. The question is whether that explanation was <em>visible at the time</em>, or only visible now in retrospect. That distinction is what tells you whether you have a process problem to fix or just a bad run of luck.</p>
<p>Categorise what you find. In practice, across many of these investigations, the patterns that surface most often are:</p>
<p><strong>Poor planning that nobody re-evaluated.</strong> The estimate was wrong at the start. The team knew mid-sprint it was wrong. Nobody raised it. The problem compounded silently until the milestone slipped.</p>
<p><strong>No upfront cross-team discussion, then a giant PR.</strong> An engineer goes heads-down, builds something substantial, opens a large merge request — and gets significant rework feedback. The time lost isn’t in the code. It’s in the weeks of parallel work that assumed alignment that was never established.</p>
<p><strong>Analysis paralysis.</strong> The overcorrection to the above. Extensive design documents, lengthy alignment meetings, decision-by-committee before any code is written. A different kind of delay, but the timeline catches it just as clearly. The pendulum swings both ways.</p>
<p><strong>Internal team silos.</strong> The most insidious, because the team often doesn’t realise it’s happening. Three weeks of frontend work sitting with one engineer while the backend engineers consider themselves done and move to the next milestone. Nobody raised a flag. Nobody re-assigned. The timeline shows a single assignee on a story for three weeks with no collaboration signals. Nobody looks negligent. They just weren’t paying attention to the right thing.</p>
<p>And here’s what makes it a leadership failure as much as a process one: the backend engineer who moved on wasn’t being careless. He thought he was being efficient. Nobody had ever told him it was acceptable — encouraged, even — to slow down and help with the frontend work, to go a bit slower individually so the team could reach the goal together. Nobody had told him to optimise for outcome, not output. He was measuring himself by what he shipped. The team was supposed to be measured by what they delivered. Those two things had quietly diverged, and the timeline is what made it visible.</p>
<p>And the thing that gets blamed loudest but is rarely the actual constraint: <strong>QA</strong>. QA teams tend to be efficient partly because they know they attract blame and compensate accordingly. If the timeline surfaces QA as a genuine bottleneck, it’s worth noting — but be prepared for it to point elsewhere first.</p>
<p><strong>External team dependencies.</strong> This one is common enough to name explicitly, because it will surface in almost every multi-team milestone — a story that couldn’t move because it was waiting on another team’s API, another team’s capacity. It shows up in the timeline as a clean gap: work picked up, nothing happened, work resumed. The team will be able to tell you exactly what they were waiting for.</p>
<p>Here’s the important thing: if this keeps appearing, it is not a problem the team can fix. It’s an organisational design problem. Teams that are structurally dependent on other teams to deliver will keep showing this pattern no matter how well they run their own process. The only durable solution is to organise around value streams — teams that own the full slice of capability they need to deliver, without hard dependencies on others for their day-to-day work. If you keep seeing external dependency gaps in your timelines, that’s the signal worth escalating.</p>
<p>The concepts of stream-aligned teams and <a href="https://less.works/less/structure/feature-teams">feature teams</a> — teams that own the full slice of capability they need to deliver, without hard dependencies on others for their day-to-day work — are two names for essentially the same idea. <a href="https://teamtopologies.com/book">Team Topologies</a> by Matthew Skelton and Manuel Pais is the clearest framework for thinking about this from an organisational design perspective. <a href="https://less.works/less/structure/feature-teams">LeSS (Large-Scale Scrum)</a> by Craig Larman and Bas Vodde arrives at the same place from a Scrum scaling angle. <a href="https://itrevolution.com/product/accelerate/">Accelerate</a> by Nicole Forsgren, Jez Humble, and Gene Kim provides the research evidence for why loosely-coupled team structures correlate directly with delivery performance. If your timelines keep surfacing the same cross-team gaps, those are the places to go next. I also explored this dynamic from a different angle in <a href="https://blog.dicko.dev/posts/the-feature-team-fallacy-why-your-agile-teams-are-just-waterfall-in-disguise/">The Feature Team Fallacy</a>.</p>
<h2>Step 6: Write the Findings Document</h2>
<p>One page. Two sections.</p>
<p><strong>What happened</strong> — the factual timeline summary, the key gaps identified, and what was happening in each one.</p>
<p><strong>What we’re changing</strong> — concrete, specific actions. Not “better communication.” Not “clearer requirements.” Things like: <em>“We will add a blocker comment in the ticket within 24 hours of being blocked, rather than surfacing it at standup.”</em> Or: <em>“Stories with external team dependencies will have that dependency explicitly tagged at planning, not discovered mid-sprint.”</em> Or: <em>“Our PR review SLA is 24 hours for first review. We will add a rotation to enforce it.”</em></p>
<p>Specificity is the test. If you can’t describe what “done” looks like for the action, it isn’t a change — it’s a wish.</p>
<h2>Who Runs This, and When</h2>
<p><strong>Who:</strong> Ideally, the engineering manager. Not because a senior engineer couldn’t do it — they often can — but because the EM is closer to a neutral third party than anyone inside the team. People have their own stake in the outcome, their own version of the sprint, their own relationships with the colleagues whose work is under the microscope. Distance matters when the findings are uncomfortable.</p>
<p>In practice, this process tends to live as tribal knowledge — something a particular EM knows how to do and runs when things go wrong, without it ever being written down or taught. That’s partly why this post exists.</p>
<p><strong>When:</strong> Not as a routine cadence. This is a diagnostic tool, not a ceremony. Run it when a milestone is missed or significantly late. Run it when a stakeholder uses the word “failed.” Run it when the same conversation keeps going in circles, when people are defending rather than analysing, when the post-retro action items look identical to last quarter’s post-retro action items. That’s the smell.</p>
<h2>The Meeting, Revisited</h2>
<p>Go back to that room. Somchai, the untouched Thai tea, the loaded silence.</p>
<p>That silence is energy. It has to go somewhere — into defensiveness, or into analysis. The difference between those two outcomes isn’t people’s intentions or their willingness to be honest. It’s whether there’s something concrete in the room to direct the energy toward.</p>
<p>The timeline is that something. Not because it exonerates anyone, and not because it assigns blame — but because it turns a competition between narratives into a shared investigation. Everyone looking at the same picture, from the same data, at the same time.</p>
<p>When you bring light to the problem, you get collaboration. Nobody wants to have this meeting again next quarter. Nobody is deliberately holding the team back. When the problem is visible and the framing is about process rather than people, everyone works to fix it.</p>
<p>The key is to be the person who walks in with the picture.</p>
<h2>The Bottom Line</h2>
<p>As Daniel Kahneman observed, <em>“We can be blind to the obvious, and we are also blind to our own blindness.”</em> Delivery failures are rarely caused by bad engineers or bad intentions. They’re caused by things that were invisible — gaps that nobody tracked, blockers that never surfaced, silos that felt like efficiency. The Execution Post-Mortem is a process for making those things visible, systematically and without theatre.</p>
<p>Your retrospective can’t do this. A retro covers two weeks of memory; a milestone failure spans months of reality. Post-it notes don’t surface a three-week frontend silo. Dot-voting doesn’t find your PR review latency.</p>
<p>What you need is the timeline. It’s unglamorous work — ticket exports, git logs, calendar cross-referencing. It takes a few hours. It requires honesty from people who’d rather move on. But it’s the only way to have the conversation that actually changes something.</p>
<p>Now if you’ll excuse me, I need to go pull a ticket history. We missed something last sprint, and I’ve already heard two compelling explanations for why. Neither of them is probably right.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Your Performance Dashboard Is Lying to You</title>
    <link>https://blog.dicko.dev/posts/your-performance-dashboard-is-lying-to-you/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/your-performance-dashboard-is-lying-to-you/</guid>
    <pubDate>Sat, 07 Mar 2026 10:06:15 GMT</pubDate>
    <category>web-performance</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>frontend-development</category>
    <category>ab-testing</category>
    <description>Or: How to Actually See the Problem Before You Try to Fix It</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/your-performance-dashboard-is-lying-to-you/cover.webp" alt=""></figure>
<p>The numbers were green. They had been green for three weeks. Somchai had pulled up the dashboard at 10:23 on a Thursday morning — Thai iced tea sweating condensation onto the vinyl floor beside his desk, the air conditioning on level 6 doing its usual indifferent job — and everything looked fine. Load time: fine. Error rate: fine. The graph lines sat flat and well-behaved, well below their thresholds. He’d closed the laptop a few minutes later and headed to planning. The user complaints had been sitting in the feedback queue the whole time.</p>
<p>This is the trap. Not that you’re not measuring. You probably are. The trap is that your dashboard is measuring the right things in the wrong way, or the wrong things entirely, and the gap between “the metrics are good” and “users are happy” is where performance work goes to die.</p>
<p>Before you can fix a performance problem, you have to be able to <em>see</em> it. That sounds obvious until you’ve spent three days optimising something your users don’t experience, or celebrated a P50 improvement while your slowest users got slower. This post is about the instruments — what they measure, what they miss, and which number from your data actually tells you something useful.</p>
<h2>Two Kinds of Measurement, Two Different Questions</h2>
<p>There are two fundamentally different approaches to measuring web performance, and confusing them is the source of a lot of wasted effort.</p>
<p><strong>Real User Monitoring (RUM)</strong> captures what actually happens when real people load your pages. Real devices, real networks, real geographic distribution, real everything. It’s the gold standard for answering “what is my users’ actual experience right now?” The limitation is that it’s retrospective — you only know about a problem after users have already experienced it. And it only works on pages you control.</p>
<p><strong>Synthetic monitoring</strong> runs automated browsers in controlled conditions — a scripted agent loads your page, measures it, and reports. It catches regressions before users see them, produces repeatable results you can track over time, and can be pointed at any URL. That last part is important: because synthetic monitoring is just “open a browser and load a page,” you can run it against your competitors just as easily as against yourself.</p>
<p>The honest summary: for measuring production performance, RUM is your primary instrument. Synthetic is how you catch regressions before they ship and how you know where you stand relative to the market. They’re complementary, not competing.</p>
<p>One important caveat: if your team runs A/B experiments — and at any serious product company, you should be — synthetic monitoring in CI becomes significantly less useful as a regression gate. Which variant do you test? The control? The treatment? Both? If a performance regression is intentional in one variant to test a hypothesis, your pipeline breaks. If you only test the control, you’re blind to what you’re actually shipping to half your users. In practice, synthetic in CI is almost useless for performance regression detection in an experimentation-heavy environment. To do it properly, your pipeline would need to know which feature flags a given PR affects, spin up synthetic runs for both the control and treatment configurations, and compare them — for every PR that touches an experiment. Almost no team does this, because almost no team has built the tooling to make it tractable. What you end up with instead is a synthetic run against your default configuration, which may bear little resemblance to what half your users are actually loading. RUM on your experiment cohorts is the only instrument that tells you whether a given variant is actually fast for real users. Don’t let a green CI pipeline give you false confidence when the thing users are experiencing is a variant your synthetic agent never loaded.</p>
<h2>The Metric Question: What Are You Actually Measuring?</h2>
<p>Here is where a lot of teams are quietly in trouble — not because they aren’t measuring, but because they’re measuring something their users don’t experience.</p>
<p>Performance metrics have a history, and it matters because the metric your team uses today is probably a product of when someone last revisited the question rather than a deliberate current decision.</p>
<h2>The browser event era</h2>
<p>For years, the standard metrics were <code>DOMContentLoaded</code> and <code>load</code> — browser events that fire as the document is parsed and resources are fetched. These made perfect sense for server-rendered pages where the content was in the initial HTML. If your page loaded, the HTML had your content.</p>
<p>Then React happened. And Angular. And Vue. And the SPA pattern spread across the industry. On a single-page application, the initial HTML is a shell. The content renders afterwards, in JavaScript, after the framework initialises. <code>DOMContentLoaded</code> fires on an empty page. <code>load</code> fires on an empty page. The metrics stay green. Users see a blank screen.</p>
<p>We hit this during our React migration. The dashboard showed no degradation. The pages were slower. Both things were simultaneously true because we were measuring the wrong moment entirely.</p>
<h2>Core Web Vitals</h2>
<p>In 2020, Google shipped <a href="https://web.dev/articles/vitals">Core Web Vitals</a>: LCP (Largest Contentful Paint), CLS (Cumulative Layout Shift), and INP (Interaction to Next Paint, which replaced FID in 2024). These metrics are backed by large-scale research on what actually correlates with how satisfied users are with a page. They’re tied to search ranking. <a href="https://www.debugbear.com/docs/rum/percentiles">P75 is the standard reporting threshold</a>. For consumer web products — ecommerce, content sites, booking flows — they’re the right starting point.</p>
<p><strong>LCP</strong> measures when the largest element on screen has painted. For a consumer page that’s usually the product image, the hero content — the thing the user came to see. LCP is a good proxy for “is this page useful yet?” when you’re building browsing experiences.</p>
<p><strong>CLS</strong> measures how much the layout shifts unexpectedly after it first appears. The number that tells you whether the page jumps around as images and ads load and knock your content out from under the user’s cursor.</p>
<p><strong>INP</strong> measures how quickly the browser shows a visible response to user input — and we’ll come back to this one in detail because it catches a pattern that most teams have in their codebase right now and don’t know about.</p>
<h2>Where Core Web Vitals don’t tell the whole story</h2>
<p>LCP measures the largest painted element. For a consumer page, that’s the product image. For a B2B data application — a pricing tool, an inventory manager, an operations queue — the largest element might be a page header or a skeleton loader. LCP fires. The data table the user actually needs hasn’t rendered yet. The user is staring at a chrome frame with no useful information. The metric is green.</p>
<p>This isn’t a flaw in Core Web Vitals. It’s a context mismatch. CWV was designed for consumer ecommerce. Applied to data-heavy B2B interfaces without adjustment, it measures the wrong moment.</p>
<p>The solution is to instrument the page yourself. We built a component — a Higher Order Component wrapper — where developers nominate the element that defines “this page is ready to use.” One wrap around the primary data table, and an event fires when it mounts. The metric captures “data the user needs is available” rather than “the largest element has painted,” which for a loading skeleton are two very different things.</p>
<p>The principle is simple: <strong>standard metrics are right for the context they were designed for.</strong> When your context differs, you extend, you don’t just accept what doesn’t fit.</p>
<p>As the statistician George Box famously observed: “All models are wrong, but some are useful.” The same is true of metrics. Use the useful ones. Know where they stop being useful.</p>
<h2>The Percentile Debate: Which Number From Your Data Actually Matters?</h2>
<p>You have your RUM data. Thousands of real user sessions, a latency distribution. Which number do you put on the dashboard?</p>
<p>This question has a right answer, a wrong answer, and a historical answer that keeps changing — which is itself the insight.</p>
<p><strong>The wrong answer: average latency.</strong> An average is a number no real user experiences. It’s distorted by outliers in both directions — a large volume of very fast requests can make a painful 95th percentile invisible, and a handful of extreme outliers can make a healthy typical experience look bad. If your performance dashboard is reporting average load time, you are flying with a broken altimeter. <a href="https://aerospike.com/blog/what-is-p99-latency/">Percentiles are the only meaningful representation of a latency distribution</a>.</p>
<p>The physicist Richard Feynman had a principle he applied to physics problems that applies here: <em>“The first principle is that you must not fool yourself — and you are the easiest person to fool.”</em> Averages fool you. Percentiles don’t.</p>
<p><strong>What each percentile actually tells you:</strong></p>
<p>Percentile What it shows The risk P50 (median) The typical user’s experience Masks the tail completely P75 Google’s CWV standard May miss genuine pain for a significant minority P90 Covers nearly everyone engineering can help Historically argued as “too noisy” P95 Near-universal experience Five percent exclusion is still large at scale P99 Right for internal service-to-service calls Dominated by noise on the public internet</p>
<p>The key insight: <strong>the right percentile is not fixed.</strong> It’s a function of how much of your tail reflects engineering problems versus environmental noise. In markets with inconsistent mobile infrastructure, the P90 tail is heavily influenced by conditions no engineering team can fix — report P90 and you’re optimising for the weather. In markets where networks have matured and device capability has standardised, the same tail is dominated by rendering and JavaScript execution problems that engineering absolutely can fix — and P75 may be letting you off too easily.</p>
<p>The question isn’t “which percentile should we use?” It’s “at what percentile does the tail stop being signal and start being noise?”</p>
<p><strong>A practical two-tier model:</strong></p>
<p>For consumer web RUM, P90 is a reasonable target if your markets are reasonably well-connected. You’re reaching nearly everyone engineering can help, without letting uncontrollable outliers drive the roadmap.</p>
<p>For internal service-to-service calls, P99 is right. You control both ends of the call. There’s no meaningful network variance between your own services. At P99, every slow request is an engineering problem, not a network one.</p>
<h2>INP: The Metric That Caught the Loading Screen Trick</h2>
<p>INP deserves its own section because it catches something common that the previous generation of metrics missed entirely.</p>
<p>FID, the metric <a href="https://web.dev/blog/inp-cwv-launch">INP replaced in March 2024</a>, measured how quickly the browser <em>started</em> handling an input — specifically, the very first user interaction on a page. INP measures how quickly the browser <em>showed a visible response</em> to the input. And it measures all interactions throughout the visit, not just the first one.</p>
<p>The practical difference: a page that swallows a click, starts an async operation, and shows a spinner 500ms later has perfect FID and terrible INP. The old metric couldn’t see this pattern. The new one can.</p>
<p>Good INP is <a href="https://web.dev/articles/inp">under 200ms at P75</a>. Needs improvement: 200–500ms. Poor: over 500ms.</p>
<p>The implementation detail most teams get wrong is the sequence. A loading indicator must appear <em>before</em> async work begins:</p>
<ul>
<li>User clicks- Button immediately shows loading state — <strong>this is the “next paint” INP measures</strong>- Async work begins
If your async operation starts before you update the DOM, INP measures the full async duration. A 500ms API call with a spinner that appears after it completes is 500ms INP, not “near-instant with a spinner.” The user clicked, saw nothing, and already clicked again.</li>
</ul>
<p>This matters beyond UX. When users don’t see a response to their input, they re-click. They interpret a non-responsive button as “my click didn’t register.” At scale, re-clicks are not just a user experience problem — they’re a reliability problem. Multiplied input events hitting your backend at unpredictable rates, particularly during degraded performance when the response takes even longer than usual, when users are <em>most</em> likely to re-click.</p>
<p>Consider a search page with mediocre INP — no loading indicators, no button state change on click. It was a known issue that hadn’t risen to priority. A query plan regression caused the underlying database query to run slower. The page got slower. With no visual feedback on click — no spinner, no disabled state, nothing — users had no way to know whether their click had registered at all. So they clicked again. And again. Each user convinced their first click had silently failed, firing the same expensive search request two, three, four times. Re-clicks multiplied the load on an already-degraded query. What would have been a slow-page incident became an outage. The root cause was the query regression. The thing that turned degraded into down was the pre-existing INP problem that had been sitting in the backlog.</p>
<p>Fix the loading indicator before it becomes the thing that turns your next slow incident into your next outage.</p>
<h2>Where This Leaves You</h2>
<p>Before you open a performance ticket, before you start profiling, before you schedule a sprint around “performance improvements,” ask three questions:</p>
<p><strong>Are you measuring what users actually experience?</strong> If you’re on a React SPA and your primary metric is <code>DOMContentLoaded</code>, the answer is no. If your B2B application&#39;s &quot;ready&quot; moment is when the data table mounts and you&#39;re measuring LCP of a skeleton loader, the answer is no.</p>
<p><strong>Are you reporting the right percentile?</strong> If your dashboard shows averages, rebuild it. If it shows P50, you’re measuring your median user while your slowest users — who are more likely to churn, more likely to complain, and sometimes more likely to be your highest-value segments in B2B — are invisible.</p>
<p><strong>Do you know what your competitors look like on the same metrics?</strong> Not to obsess over them, but because there is no absolute threshold at which performance “passes.” There’s only faster or slower than the alternatives your users have. Synthetic monitoring against a few benchmark URLs costs almost nothing and tells you whether your users are tolerating you or choosing you.</p>
<p>As Einstein put it — with characteristic economy — <em>“The important thing is to not stop questioning.”</em> Your monitoring stack is only as good as your willingness to interrogate it. RUM, synthetic, the right metric for the right surface, the right percentile for the right context. None of them works alone. Together, they tell you what’s actually happening — but only if you keep asking whether what you’re seeing is the whole picture.</p>
<p>The dashboard that says everything is fine while users are complaining isn’t a mystery. It’s a measurement problem. And measurement problems are the most fixable kind — once you can see what’s wrong, you can fix it. The hard part is admitting that the green numbers might not mean what you think they mean.</p>
<p>In the next post, we’ll look at where performance problems actually hide in a frontend codebase — because once you can see the metrics clearly, the next question is what to do about them.</p>
<p><em>This is the first in a three-part series on web performance engineering. Post 2 covers where performance problems actually live in the frontend. Post 3 covers how to build a performance culture that sticks.</em></p>
]]></content:encoded>
  </item>
  <item>
    <title>A Brief History of Web Performance at Agoda</title>
    <link>https://blog.dicko.dev/posts/a-brief-history-of-web-performance-at-agoda/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/a-brief-history-of-web-performance-at-agoda/</guid>
    <pubDate>Sat, 07 Mar 2026 06:58:06 GMT</pubDate>
    <category>web-performance</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>core-web-vitals</category>
    <category>tech</category>
    <description>Or: What Ten Years of Measuring the Wrong Things Taught Us About Measuring the Right Ones</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/a-brief-history-of-web-performance-at-agoda/cover.webp" alt=""></figure>
<p>The dashboard was green. Every metric we cared about was sitting comfortably within target. It was a Tuesday morning on level 6, the air conditioning was losing its daily battle with the Bangkok heat, and someone had made fresh Thai tea in the kitchen — you could smell the condensed milk from three desks away. I was looking at the numbers with the quiet satisfaction of someone who genuinely believed they understood what was happening in production.</p>
<p>We didn’t. The users were staring at an empty React shell while our DOM load metric cheerfully reported that everything was fine. The measurement infrastructure we’d built and trusted for years had silently stopped measuring anything useful, and the numbers had stayed green the whole time.</p>
<p>That’s the honest version of how this story starts.</p>
<p>What follows isn’t a “here’s how we do performance” post. Those posts describe a tidy current state — the architecture diagram, the metrics that matter, the tools you should adopt. This is the other kind: the version with the abandoned experiments, the wrong turns, the metrics we built before the industry had names for them, and the debates that ran twice and reached opposite conclusions. It’s a history of how a team figured out, iteratively and sometimes painfully, what “fast” actually means for users — and how to measure it honestly.</p>
<p>If some of this sounds familiar, that’s the point.</p>
<h2>Phase 1: DOM Load and Boomerang (2016)</h2>
<p>Agoda’s first serious investment in Real User Monitoring came in 2016, built around a fork of <a href="https://www.oreilly.com/library/view/high-performance-web/9780596529307/"><strong>Yahoo’s Boomerang library</strong></a> — one of the foundational open-source RUM tools, a direct descendant of Steve Souders’ pioneering work on web performance in the mid-2000s. At the time, Agoda was server-side rendered. Boomerang injected a lightweight JavaScript agent into every page, collected browser timing events, and fed the data into Agoda’s internal analytics platform.</p>
<p>The primary metric was <strong>DOM load event time</strong>. For a server-rendered page, this is a reasonable proxy for “the page is usable.” The server sends HTML, the browser parses it, the DOM loads, and the content is there. Simple, defensible, and — as long as your architecture doesn’t change — correct.</p>
<p>The integration that proved most valuable wasn’t the metric itself but where the data landed. Performance data flowed into Agoda’s A/B testing infrastructure from the start. Every experiment automatically generated a performance comparison between variants. That decision — treating performance as a first-class output of every test, not a separate concern — shaped everything that came after.</p>
<p><strong>The percentile debate, first iteration</strong></p>
<p>When you have real user data across thousands of sessions you immediately face the question: which number do you actually report? Median? 75th percentile? 90th?</p>
<p>Agoda settled on <strong>P90</strong> as the primary optimisation target, and the reasoning was worth understanding. At Agoda’s traffic scale, even a small percentage improvement on the slower end of your distribution represents a large absolute number of users. The tail at P90 — the slowest tenth of your users — is also the group most likely to abandon. Improving their experience can deliver more commercial impact than shaving milliseconds from the median.</p>
<p>The counterargument at the time was that P90, on mid-2010s mobile networks and hardware in Asia, was dominated by noise — bad connections, underpowered devices, unusual browser states. Optimising for P90 was, some argued, optimising for chaos rather than for real user behaviour. There was a point to that. In practice, Agoda split: P75 for SLAs, P90 and P50 as additional lenses for experimentation. The debate wasn’t resolved so much as it was made productive.</p>
<h2>Phase 2: Synthetic Monitoring and the Competitor Charts (2017)</h2>
<p>In 2017, Agoda deployed a fleet of synthetic monitoring agents in offices worldwide. The goal was appealing: approximate real geographic user experience without waiting for RUM data to accumulate, and catch regressions before users ever saw them.</p>
<p>The operational reality was a maintenance problem. Agents needed regular restarts. Power outages in remote offices took monitoring nodes offline without obvious signals. State accumulated and required periodic reloading. Running a global fleet of monitoring machines is infrastructure work — and it doesn’t naturally belong to a product engineering team. Eventually the experiment was abandoned.</p>
<p>But synthetic monitoring did one thing that RUM fundamentally cannot: <strong>it could load any URL, including competitors’.</strong></p>
<p>Agoda pointed the agents at Booking.com, Expedia, and Airbnb alongside internal pages. Every bi-weekly performance review, engineers saw Agoda’s load time plotted on the same chart as the competition. The question stopped being “did we hit our target” and became “are we faster than Airbnb yet?”</p>
<p>Robert Cialdini, whose work on influence and social behaviour remains some of the most practically useful psychology written in the last fifty years, found that people are far more motivated by knowing where they stand relative to others than by knowing where they stand relative to a target. Abstract millisecond improvements are hard to make feel meaningful. A race against a named competitor is not. Engineers knew exactly who they were chasing and roughly where the gap was. Goals had faces.</p>
<p>We caught Booking.com on mobile load time for a period. Airbnb stayed ahead of us longer than we’d like to admit — they were genuinely good at this. But the charts kept people honest in ways that internal targets rarely do.</p>
<p>The synthetic fleet is gone. The insight isn’t. Tools like <a href="https://www.webpagetest.org/">WebPageTest</a> let you run competitor benchmarks today without maintaining your own agent infrastructure — the fragile part of the experiment turned out to be separable from the valuable part.</p>
<h2>Phase 3: React, the DOM Load Problem, and Critical Mark (2018)</h2>
<p>When Agoda migrated to React, DOM load time stopped being useful as a primary metric. The initial HTML document is a shell — it loads quickly, because there’s nothing in it. JavaScript bundles then load, parse, and execute. React renders. Data fetches complete. The content the user actually came for appears. DOM load fires long before any of this is done.</p>
<p>The metric looked excellent. Users were waiting.</p>
<p>This is the gap that browser timing events don’t cover: the difference between “the page shell loaded” and “the user can start working.” In 2018 — two years before Google launched Core Web Vitals — Agoda built a custom metric to fill it.</p>
<p>We called it the <strong>Critical Mark</strong>.</p>
<p>The implementation was deliberately simple: a Higher-Order Component that injects a callback prop. Engineers nominate one component in each page’s tree as the readiness point — typically the primary data payload, the hotel list, the search results, the booking summary. When that component mounts, a RUM timing event fires and lands in the same infrastructure as every other performance measurement.</p>
<p>One HOC wrap. One nomination per page type. A metric that answers the question users actually care about.</p>
<p>The data connected directly to A/B testing, as DOM load data had before. Every test continued to produce a performance comparison. The metric predated the industry equivalent, was built to solve a specific gap, and quietly became load-bearing for how Agoda thought about performance.</p>
<h2>Phase 4: Google Core Web Vitals — Adoption and Adaptation (post-2020)</h2>
<p>Google launched Core Web Vitals in 2020: Largest Contentful Paint, Cumulative Layout Shift, and First Input Delay. For most web products, LCP does something very similar to what Critical Mark was doing — measuring when the primary content element is visible to the user.</p>
<p>Agoda evaluated the <code>web-vitals.js</code> library and adopted it, but not off the shelf.</p>
<p>The fork added two things. First, internal telemetry routing: Web Vitals events flow into Agoda’s own data platform, sitting alongside product analytics and A/B testing data in the same system. Second, per-variant performance comparison: every experiment now automatically outputs Web Vitals comparisons alongside product metrics. This, again, is the most reliable way to know whether a change genuinely improved performance — not synthetic tests, not CI analysis, but real users split across controlled variants experiencing the actual product.</p>
<p>At this point the original Critical Mark was retired for the consumer product. LCP, and later <a href="https://web.dev/blog/inp-cwv-launch">INP (which replaced FID in 2024 as a Core Web Vital)</a>, covered consumer ecommerce well enough. The custom metric had been superseded — at least for that context.</p>
<p><strong>The percentile debate, second iteration</strong></p>
<p>The same question came back around, and this time it inverted. The argument in 2016 had been “P90 is too aggressive, the tail is too noisy.” By the early 2020s, the network and hardware landscape had shifted. Mobile connections were better. The P90 tail was less dominated by chaos and more representative of real users with real devices on real networks who could genuinely benefit from engineering work.</p>
<p>The question had flipped: not “should we lower the target” but “should we go further than Google’s P75 recommendation?” Agoda stayed at P75 as the primary SLA, but P90 continued as an active lens in some teams. The answer depends on your context. The debate being worth having at all says something useful about how much the landscape had changed.</p>
<h2>Phase 5: The B2B Extranet Project and Critical Mark Returns (2022)</h2>
<p>In 2022, a dedicated performance initiative launched for Agoda’s B2B extranet — the platform hotel partners use to manage inventory, pricing, and promotions. The scope was deliberately cultural as well as technical, and that wasn’t an accident — it came directly from what the frontend performance team had learned working through all of the phases above on the main consumer funnel.</p>
<p>The lesson they’d arrived at the hard way: you cannot have one team that owns performance, fixes performance, or even improves performance in isolation. If no one else cares about it, everyone else makes it worse, and you end up with a single team perpetually trying to fix the mistakes of every other team shipping code. Worse, the moment you have a dedicated performance team, people tend to develop the attitude of “don’t we have a team for that? why is it my problem?” The performance team needs to be an enabler and an educator: helping teams with tooling, setting SLAs, understanding how the DOM renders, explaining why this particular thing is slow. The knowledge has to spread, or it doesn’t hold.</p>
<p>The 2022 extranet project was built on that foundation. The team recognised that <strong>performance isn’t a team, it’s a culture</strong>, and the deliverable wasn’t just tooling — it was rolling out norms and practices to the engineering teams building extranet products.</p>
<p>Working on the extranet exposed a familiar gap. LCP fires when the page header and shell load, not when the data hotel partners actually need appears. A revenue manager opening a pricing page gets a “loaded” LCP signal while the rate table is still fetching. The metric the industry had standardised on was measuring the wrong thing for this context.</p>
<p>We began so see patterns like this:</p>
<figure><img src="https://blog.dicko.dev/posts/a-brief-history-of-web-performance-at-agoda/image_2.webp" alt="" loading="lazy"></figure>
<p>You can see on the left at 1800ms the LCP mark is hit, but the data isn’t there, the data comes almost a full second later.</p>
<p>And when we applied this to very data heavy complex pages, it’s even worse.</p>
<figure><img src="https://blog.dicko.dev/posts/a-brief-history-of-web-performance-at-agoda/image_3.webp" alt="" loading="lazy"></figure>
<p>The solution was to reach back for Critical Mark. Same pattern, same HOC approach, applied to B2B pages where the primary content is a data table rather than a media element. It wasn’t rediscovered — it was deliberately retrieved.</p>
<p>There’s something worth noting in that arc. We built a metric before the industry had an equivalent, the industry eventually caught up, the industry standard superseded our custom metric for consumer ecommerce, and then we needed the custom metric again for a context the industry standard wasn’t designed for. LCP and Critical Mark now run alongside each other: LCP for tracking bundle optimisation impact and early page load characteristics, Critical Mark for measuring what hotel partners experience as “the page is ready.”</p>
<h2>Phase 6: What We’re Exploring Now (2024–present)</h2>
<p>Two things are currently in experimentation. Neither is in production. Both might fail, failure is always an option, but fail with learnings.</p>
<p><strong>Compression Dictionary Transport (CDT)</strong> reduces the size of repeated JavaScript bundles by using a shared dictionary for compression. POCs are complete. Before committing to it, the team went looking for evidence of the problem it solves — a visible signal in LCP, Critical Mark, or other existing metrics that bundle size was meaningfully hurting users. So far that signal hasn’t appeared clearly in the data. The team’s position is straightforward: find the problem before deploying the solution. CDT is a validated technology. Deploying it without evidence that you have the problem it solves is optimising into the unknown.</p>
<p><a href="https://docs.zephyr-cloud.io/"><strong>Zephyr Cloud</strong></a> <strong>edge tree shaking</strong> addresses a different structural issue. Agoda runs a significant number of live experiment variants simultaneously. JavaScript bundles carry code for all active variants — code the current user will never execute, because they’re assigned to only one variant of each experiment. Edge tree shaking strips dead-variant code at serve time based on the user’s experiment assignments. Agoda is Zephyr Cloud’s first on-premises customer; at the scale of infrastructure involved, the SaaS hosting model doesn’t apply. Zephyr calls it “Bring Your Own Infrastructure” — the on-prem deployment model runs on Agoda’s own infrastructure. The module federation underpinnings connect back to <a href="https://blog.dicko.dev/posts/micro-frontend-strategy-choosing-between-shell-apps-and-independent-systems/">how we think about micro-frontend architecture</a> more broadly.</p>
<p>These might work. They might not. We’ll test them properly and find out.</p>
<h2>What the History Actually Teaches</h2>
<p>If there’s a thread running through ten years of this, it’s something the physician and writer Atul Gawande captured well when writing about how medicine learns: “We have not asked surgeons to have the same results every time. We have asked them to have better results over time.”</p>
<p>That’s the honest version of what a performance programme looks like. Not a moment when you get it right. A series of iterations where you get it less wrong — because the technology changes, the architecture changes, the user base changes, and the metrics that were correct last year measure something irrelevant today.</p>
<p>DOM load was right until React made it wrong. Core Web Vitals were right for consumer ecommerce until they were wrong for B2B extranet. P90 was too aggressive until the networks improved and it wasn’t anymore.</p>
<p>The measurement infrastructure we trust today will need revisiting. That’s not a failure of the approach. It’s the approach working exactly as it should — <a href="https://blog.dicko.dev/posts/semantic-monitoring-the-question-youre-not-asking-about-your-production-systems/">measuring what users actually experience, not what servers report</a>.</p>
<p>Now, if you’ll excuse me, I need to go check whether our Critical Mark nominations on the extranet are still pointing at the right components. I’m fairly certain one of them hasn’t been updated since the last redesign, which would mean we’ve been measuring the wrong thing confidently for several months.</p>
<p>Apparently some problems are perennial.</p>
]]></content:encoded>
  </item>
  <item>
    <title>When Designers and Engineers Disagree (And Both Are Right)</title>
    <link>https://blog.dicko.dev/posts/when-designers-and-engineers-disagree-and-both-are-right/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/when-designers-and-engineers-disagree-and-both-are-right/</guid>
    <pubDate>Wed, 04 Mar 2026 14:48:53 GMT</pubDate>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>design-systems</category>
    <category>ux-design</category>
    <category>product-engineering</category>
    <description>Or: Why the Best Products Come from Arguments Nobody Wins</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/when-designers-and-engineers-disagree-and-both-are-right/cover.webp" alt=""></figure>
<p>The Slack thread had forty-seven replies. Nai, one of our designers, had stopped typing and started walking — weaving through the open-plan desks on level 6 toward the engineering area, which is never a good sign. The air conditioning hummed the way it always does when a conversation is about to get interesting: just loud enough that people at neighbouring desks can pretend they’re not listening. Someone’s Thai iced tea sat untouched on a desk, condensation slowly working its way down the plastic cup.</p>
<p>This wasn’t a disagreement about whether something looked nice. This was two people, both excellent at their jobs, both absolutely certain they were protecting the user — and both right. That moment, the one where a designer and an engineer lock eyes across a design review and neither blinks, is the moment that determines whether your product team is genuinely collaborative or just politely dysfunctional.</p>
<p>As the architect Charles Eames once said, “The details are not the details. They make the design.” He understood something most product teams don’t: details matter enormously, but <em>which</em> details matter is the real question. And when a designer and an engineer disagree, they’re almost never arguing about whether details matter — they’re arguing about <em>which detail matters more</em>.</p>
<h2>The Header That Started a War</h2>
<p>Here’s the story. Our design team wanted the site header to change its background colour depending on whether the user was on a modern page or a legacy one. The logic was sound from a design perspective: each page has its own visual identity, its own mood, and the header should feel like it <em>belongs</em> to the page you’re on. Mismatched headers create a subtle but real sense that the product is stitched together from different eras — which, to be fair, it was.</p>
<p>Engineering pushed back. The argument was equally sound, but rooted in a completely different discipline: web performance. Specifically, Cumulative Layout Shift and contentful paint metrics. When you navigate between pages, you expect the header — the single most persistent element on the screen — to remain stable. A header that changes colour on navigation isn’t just a visual inconsistency; it’s a layout signal that tells you something fundamental has changed. Google’s Core Web Vitals penalise exactly this kind of visual instability. And users, even when they can’t articulate it, perceive a page where the chrome stays consistent as faster and more trustworthy than one where elements shift around.</p>
<p>Both sides were right. The designers were optimising for visual coherence <em>within</em> a page. The engineers were optimising for perceptual stability <em>across</em> pages. Neither concern was trivial. Neither was wrong.</p>
<p>How did it end? The header stayed consistent. Engineering won this one — partly because the CLS argument had measurable data behind it, and partly, if I’m being honest, because I was stubborn about it. That honesty matters: not every design-engineering disagreement resolves through elegant compromise. Sometimes one side wins. The question is whether the winning argument was principled and whether the losing side was heard.</p>
<p>But here’s the thing worth sitting with: did the right side win? Probably. The header consistency makes the product feel faster and more stable as users navigate. The designers’ concern about legacy pages looking stitched together was real — but the answer turned out to be <em>fixing the legacy pages</em>, not making the header pretend they didn’t exist. Sometimes the right resolution to a design-engineering disagreement isn’t a compromise at all — it’s recognising that both sides identified a real problem, but only one side proposed a solution to the right problem.</p>
<h2>Why We Talk Past Each Other</h2>
<p>You’ve been in these conversations. You know the dynamic. The designer presents a vision that makes your stomach drop because you can already see the fourteen edge cases it doesn’t account for. Or you’re the one watching an engineer casually propose “just a simple dropdown” that will absolutely destroy the interaction pattern your team spent three weeks refining.</p>
<p>The pattern is always the same. Designers protect the user’s experience. Engineers protect the system’s health. And the product suffers when either side wins absolutely.</p>
<p>What designers see that engineers often miss: spacing that feels wrong isn’t aesthetic fussiness — visual rhythm affects comprehension and trust. That animation isn’t decoration — it’s communication, guiding attention and signalling state changes. The ugly error state isn’t a nice-to-have — a beautiful happy path with a broken edge case teaches users your product is fragile. Inconsistency across components isn’t minor — every deviation is a micro-decision the user has to make.</p>
<p>What engineers see that designers often miss: what works for a hundred items shatters at ten thousand — the demo isn’t the product and doesn’t scale to the 99th percentile that we measure the server side response time sla on. Every custom component is something someone has to maintain, test, and keep compatible indefinitely. That beautiful unified experience might require coupling three services that currently deploy independently. And design tools don’t have network latency, memory constraints, or accessibility requirements baked in.</p>
<p>As the physicist Richard Feynman put it, “The first principle is that you must not fool yourself — and you are the easiest person to fool.” Both designers and engineers are susceptible to this. Designers can fool themselves that fidelity matters more than feasibility. Engineers can fool themselves that technical constraints are more rigid than they actually are. The magic happens when both sides are honest about what they’re actually optimising for — and whether that’s the right thing to optimise for right now.</p>
<h2>Five Hundred Shades of Blue</h2>
<p>A decade ago, Agoda’s design-engineering interface was a classic one-way handoff. Designers produced mockups. Engineers figured out what CSS to write. Nobody checked whether the blue in this designer’s mockup was the same blue as that designer’s mockup.</p>
<p>The result: at one point, the Agoda website had something like five hundred different shades of blue. Not because anyone was incompetent — because the system had no mechanism for consistency. Every designer was making independent colour choices. Every engineer was faithfully implementing those independent choices. The cumulative result was a visual experience that looked like it was designed by a committee that never met.</p>
<p>This is the natural outcome of one-way handoffs to engineers: code <a href="https://blog.dicko.dev/posts/code-entropy-the-silent-killer-of-engineering-velocity/">entropy</a>. When anyone can add a shade of blue but nobody owns the palette, inconsistency isn’t a bug — it’s the inevitable thermodynamic outcome.</p>
<p>The first attempt at a design system was open to contributions from anyone. This sounds collaborative, but in practice, it succumbed to a similar <a href="https://blog.dicko.dev/posts/code-entropy-the-silent-killer-of-engineering-velocity/">entropy </a>it was meant to prevent. Without governance, a democratised design system is just a more organised version of the chaos it replaced. Components proliferated. Variants multiplied. The system became so permissive that it stopped being a system. I’ve rarely seen collective code ownership work, too many owners always leads to no ownership at some tipping point (actual result may vary on this).</p>
<p>As the management thinker W. Edwards Deming observed, “A bad system will beat a good person every time.” We had good people in a bad system. The fix wasn’t better people — it was better infrastructure.</p>
<h2>The Design System as Shared Infrastructure</h2>
<p>What works today is <a href="https://medium.com/agoda-engineering/pixels-to-product-innovating-with-design-system-1062f9645a34">ADS — Agoda’s Design System</a> — and the key insight is in the team structure, not the component library. A group of engineers and designers worked together to raise the five-hundred-shades-of-blue problem as a systemic issue and propose a fundamentally different way of working. The result was a dedicated ADS team that <em>builds and owns</em> the design system.</p>
<p>Here’s the common mistake most companies make: treating the design system as a <em>design</em> thing that engineers adopt. This creates two sources of truth — a Figma component library and a coded component library — that inevitably drift apart. The design system becomes the very source of the inconsistency it was meant to prevent.</p>
<p>ADS is owned by a central team that comprises both designers and engineers. This isn’t a detail — it’s the whole game. When the design system team includes people who think in pixels <em>and</em> people who think in code, the system naturally evolves as shared infrastructure rather than a design artefact that engineering has to consume.</p>
<p>What this looks like in practice: designers compose pages from ADS components in Figma. Engineers compose pages from the same ADS components in React. Because both sides use the same building blocks, “pixel perfect” isn’t something engineers aim for — it’s something they get automatically by using the right component. The team uses Playwright component testing for visual regression — automated tests that catch when a rendered component drifts from its expected appearance. There are Figma-to-code parity checks that verify the Figma components and their coded counterparts stay in sync.</p>
<p>And then there’s the part that still slightly amazes me. We’ve built an MCP — Model Context Protocol — integration for ADS, which means coding agents can reference the design system directly. The actual workflow today often looks like this: an engineer takes a screenshot of the Figma design, pastes it into Cursor, and asks it to generate the React code. Cursor uses the ADS MCP to look up the correct components, their props, their variants — and generates code that uses the real design system components. The Figma MCP itself is still rubbish, so the screenshot-to-Cursor-to-MCP pipeline is the pragmatic workaround. First-pass success rate is around 40–50% — not perfect, but even when it’s wrong, it’s usually <em>almost</em> right. A few manual tweaks and you’re done. That’s a significant time saving over writing component composition from scratch.</p>
<p>Brad Frost — the creator of Atomic Design — has argued that “handoff” and “automation” are the two most damaging buzzwords in design systems, because they strip away the humanity and collaboration that actually make good products. Our model validates his point. We didn’t solve the handoff problem by automating it. We solved it by making the handoff unnecessary.</p>
<h2>Design Is How Engineers Understand the Problem</h2>
<p>Here’s where this connects to something bigger. Most “design-engineering collaboration” articles frame the relationship as a workflow problem — how do we move artefacts from one discipline to another with less friction? That framing is wrong. Design is how engineers understand the problem they’re solving.</p>
<p>In a Product Engineering model, everything is a problem you’re solving. The engineer’s job isn’t to implement a spec — it’s to solve a user problem using technology. But you can’t solve a problem you don’t understand. And a key understanding of user problems lives in the design discipline — in user research, in Jobs-to-be-Done analysis, in the empathy that comes from watching someone struggle with your product.</p>
<p>At Agoda, engineers get involved in design thinking exercises and the Jobs-to-be-Done framework. This isn’t about making engineers into designers. It’s about three things.</p>
<p>First, business acumen. When an engineer watches a user try to complete a booking and struggle with a particular flow, they understand the business impact viscerally. That’s different from reading a ticket that says “improve booking flow conversion.”</p>
<p>Second, empathy with the product. Engineers who understand why a design decision was made are better equipped to preserve the intent during implementation — and to push back when the intent is unclear or the design doesn’t actually serve the user.</p>
<p>Third, better engineering decisions. An engineer who understands the user’s job-to-be-done makes different technical choices than one who’s implementing a spec. They anticipate edge cases the designer didn’t think of. They suggest simpler solutions that achieve the same outcome. They know which corners to cut and which are load-bearing.</p>
<p>This is why the header story is instructive. If the designer and engineer had only argued about visual consistency versus CLS scores, it’s a stalemate. But when both sides ask “what problem are we solving for the user?”, the answer gets richer: users need to feel that the product is coherent <em>and</em> that it’s fast and stable. That’s not a design problem or an engineering problem — it’s a product problem that both disciplines illuminate from different angles.</p>
<h2>Tactics for When You’re in the Room and It’s Getting Heated</h2>
<p>You’re going to find yourself in these disagreements. Here’s what works.</p>
<p>Start with “What are we optimising for?” Before debating the solution, agree on the goal. Are we optimising for launch speed, user delight, system simplicity, or long-term maintainability? Often, the disagreement dissolves when both sides realise they’re optimising for different things — and neither was explicitly stated.</p>
<p>Quantify the trade-off. Put numbers on it. “This custom animation adds two days of engineering and one day of QA. Is that trade-off worth it for this feature?” Sometimes yes, sometimes no — but making it explicit prevents resentment.</p>
<p>Explore the spectrum, not the binary. Disagreements often present as “your design” versus “my technical constraint.” The productive move is to map the spectrum. “Here’s the full version — two weeks. Here’s a good-enough version — three days. Here’s the minimum — one day. Which point gives us the best value for the investment?”</p>
<p>Timebox the disagreement. If you can’t resolve it in thirty minutes, you don’t have enough information. The next step should be research — user testing, a technical spike, data analysis — not more arguing.</p>
<p>And the one I use most: the “two weeks from now” test. If we ship this version, will anyone — user, designer, engineer, stakeholder — be upset about this trade-off in two weeks? If not, ship it. If yes, invest more.</p>
<h2>The Bottom Line</h2>
<p>Design-engineering tension isn’t a bug in your process. It’s the feature that produces your best work — if you have the organisational muscle to use it. Kill the handoff. Build shared infrastructure. Get engineers into user research and designers into architecture discussions. Not to approve each other’s work, but to understand it. Celebrate when a designer changes direction based on engineering feedback. Celebrate when an engineer goes beyond the spec to get the details right. Build the buddy system that gives you both consistency and collaboration.</p>
<p>And when the disagreement hits — and it will — remember that you’re not arguing about who’s right. You’re arguing about which detail matters more. The answer is almost always “both, and here’s the third option neither of us saw yet.”</p>
<p>Now, if you’ll excuse me, I need to go review a Figma file. Apparently the designer buddy for one of my teams has forty-seven comments, and I have a feeling I know how this ends.</p>
]]></content:encoded>
  </item>
  <item>
    <title>The Business Logic Your Engineers Can’t See</title>
    <link>https://blog.dicko.dev/posts/the-business-logic-your-engineers-cant-see/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/the-business-logic-your-engineers-cant-see/</guid>
    <pubDate>Tue, 03 Mar 2026 13:07:52 GMT</pubDate>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>domain-driven-design</category>
    <category>engineering-leadership</category>
    <category>business-acumen</category>
    <description>Or: Why Technical Excellence Without Domain Understanding Is Expensive</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/the-business-logic-your-engineers-cant-see/cover.webp" alt=""></figure>
<p>It was 2:15 on a Wednesday on level 6, and the aircon was losing its quiet war against a room with too many bodies and too many monitors. Someone had made Thai iced tea in the kitchen, and that sweet burnt-sugar smell drifted across the open plan. Somchai had his hand up again. Not the confident hand of someone with an answer — the tentative half-raise of someone about to ask the PO what colour a button should be. Three seats away, Namfon had the same Slack DM open, cursor blinking, the same question queued behind two others. The junior PO sat with her shoulders up near her ears, backlog untouched on her second monitor, a half-eaten bag of Lays slowly going stale beside her keyboard. She wasn’t thinking about product strategy. She was answering her fourteenth clarification question of the day.</p>
<p>That team had something most engineering organisations would kill for: a dedicated Product Owner with nothing else to do but support them. And it produced the worst domain understanding of any team in the org.</p>
<p>As the management theorist Russell Ackoff once put it, “The righter you do the wrong thing, the wronger you become.” We’ve spent decades getting better at building software — better tools, better frameworks, better processes — while consistently underinvesting in the thing that determines whether the software we build actually matters: understanding the business it serves.</p>
<p>Here’s the dirty secret of software engineering education: we teach people how to code, how to architect systems, how to write tests — and then drop them into a domain they know nothing about and wonder why the software doesn’t match the business. The gap between “technically correct” and “solves the actual problem” is almost always a gap in business understanding, not a gap in technical skill.</p>
<p>Most engineering failures aren’t engineering failures. They’re comprehension failures. The code compiles. The tests pass. The service scales. And the feature does something nobody asked for, because the engineer who built it didn’t understand the business domain deeply enough to know why the requirement existed in the first place.</p>
<h2>The PO Who Answered Every Question (And Why That Was the Problem)</h2>
<p>Most posts about engineers and business understanding open with a disaster: the pricing rule nobody understood, the regulation nobody checked, the feature that missed the market. Those stories are real, but they’re not the interesting failure mode. The interesting failure mode is the one that looks like success.</p>
<p>That team I mentioned — they shipped working software. Nothing broke. Every ticket was completed, every acceptance criterion met. If you looked at velocity charts, they were a model squad. But beneath the surface, something corrosive was happening.</p>
<p>The engineers had stopped investing in understanding <em>why</em> things worked the way they did. They didn’t need to — they could ask. Every question became a lookup: “what should this do?” instead of “why does this exist?” They built features correctly but couldn’t connect what they’d built to the business reason behind it. When experiments showed bias, the engineers couldn’t spot it — because they had no mental model of how users actually used the product. They had no map of the user journey, just a series of tickets with instant answers attached.</p>
<p>The junior PO, meanwhile, was drowning. She’d been hired as a strategic partner and had become an on-demand FAQ service. Her product thinking atrophied because every minute was consumed by reactive clarification instead of proactive planning.</p>
<blockquote>
<p>“The heart of software is its ability to solve domain-related problems for its user.”* — Eric Evans, *Domain-Driven Design</p>
</blockquote>
<p>The paradox: the team with the most access to domain knowledge had the least domain understanding. Not because they were lazy. Because the system they were in removed the incentive to learn.</p>
<p>One of my solutions was counterintuitive — give engineers <em>less</em> PO time, not more. Tell the PO it’s OK not to come to every standup. The constraint of limited access forces engineers to invest in understanding the domain upfront, because they can’t rely on just-in-time clarification for every decision. When your PO is in meetings and unavailable — which is a common occurrence at Agoda — you learn to ask deeper questions during planning so you don’t get caught out later. That friction isn’t a bug. It’s the mechanism that builds domain fluency.</p>
<h2>Why Engineers Don’t Understand the Business (And Why It’s Not Their Fault)</h2>
<p>Let me be clear: this isn’t about lazy or incurious engineers. I’ve never met an engineer who <em>wanted</em> to build the wrong thing. This is about structural barriers that organisations erect — usually without realising it — between engineers and the business knowledge they need.</p>
<p><strong>Tickets strip context — and “Definition of Ready” makes it worse.</strong> Jira stories reduce complex business rules to acceptance criteria. The <em>why</em> gets lost somewhere between the product meeting and the ticket description. By the time an engineer reads “as a user, I want to see the adjusted price including commission,” the entire history of why that commission structure exists, which partner contracts it reflects, and what regulatory constraints shaped it has evaporated.</p>
<p>I had a team that tried to solve this by creating a “Definition of Ready” — a checklist of everything a Jira story needed before engineers would accept it into a sprint. On the surface, it sounds reasonable. In practice, it was a red flag. The Definition of Ready pushed all the context-gathering work onto the PO and turned engineers into passive consumers of pre-digested requirements. The PO did the legwork, the engineers reviewed the output, and nobody on the engineering side built any understanding of the domain in the process. The ticket arrived “ready” — neat, complete, and completely disconnected from the messy business reality that shaped it. Engineers should be <em>part</em> of that legwork. The act of helping refine a requirement — talking to stakeholders, understanding the constraints, asking why the rule exists — is how domain knowledge transfers. When you outsource that to a checklist, you get well-formatted tickets and engineers who still can’t tell you <em>why</em> the business needed what they just built.</p>
<p><strong>Specialisation creates tunnel vision.</strong> In microservices organisations, an engineer might work on “the payment service” for two years and never understand how payments fit into the wider booking flow. They know everything about authorisation and capture, nothing about why the dependency graph of the booking systems upstream needs to route through their service in a particular order.</p>
<p><strong>Domain experts are busy — or too available.</strong> This one cuts both ways. When POs and business analysts are stretched thin, they answer questions when asked but don’t proactively teach the domain. And as we saw with our always-available PO, when they <em>are</em> available, engineers outsource understanding instead of building it. The sweet spot is somewhere in the middle, and most organisations land at one extreme or the other.</p>
<p><strong>Onboarding is purely technical.</strong> Your new hire bootcamp probably covers the tech stack, the CI/CD pipeline, and the coding standards. Does it cover the business model? The key business processes? The regulatory environment? If a new engineer joins the pricing team, do they learn how OTA pricing actually works — the commission structures, the partner relationships, the currency conversion chains — or do they just learn the codebase? If you’re only teaching the codebase, you’re setting engineers up to build the wrong thing.</p>
<p><strong>Domain knowledge is tacit.</strong> The people who really understand the business carry it in their heads. It’s not documented because it changes constantly and documentation goes stale before the ink dries. The result is tribal knowledge — “you need to talk to Namfon about that service, she’s the only one who knows why it does that thing with Japanese tax calculations” — which is fragile, unscalable, and a single point of failure.</p>
<p>The pattern across all of these: domain knowledge is treated as someone else’s job, but the consequences of its absence land squarely on engineering. The engineer who doesn’t understand why a rule exists can’t flag when a new feature contradicts it. They can’t simplify the implementation because they don’t know which constraints are real and which are accidents of history. They can’t push back on bad requirements because they don’t have the standing that comes from domain fluency.</p>
<h2>The Sixty-Node Booking Graph (Or: When “Simple” Is a Domain Question)</h2>
<p>Let me make this concrete with an example from our world.</p>
<p>Agoda’s booking engine has a dependency graph of roughly 60 nodes that orchestrates a single booking. If you’re an engineer who’s never worked in travel tech, your first reaction is probably: “That’s insane. Why is it so complex? Who over-engineered this?”</p>
<p>That reaction is a domain comprehension failure. Here’s why.</p>
<p><strong>The source of truth isn’t you.</strong> Agoda sells accommodation from global hotel chains, wholesale travel suppliers, all the way down to individual homes (think Airbnb operators). For many of these, the source of truth for “is this the right price?” or “is this room still available?” isn’t Agoda — it’s the supplier. A chain hotel might respond in milliseconds via a standardised API. A small guesthouse in rural Thailand might be on a channel manager that takes seconds to confirm. The dependency graph has to accommodate real-time verification against external systems with wildly different APIs, response times, and reliability characteristics — with fallbacks, timeouts, and reconciliation logic for each supplier type.</p>
<p><strong>Three hundred payment methods across dozens of currencies.</strong> Agoda supports over 300 payment methods in a massive range of currencies. This isn’t a design choice — it’s a business requirement driven by the reality of Asian markets. GrabPay in Southeast Asia, Alipay and WeChat Pay in China, bank transfers in Japan, UPI in India — each with their own authentication flows, settlement timelines, currency handling, and failure modes. A dependency graph that only supported credit cards would be a fraction of the complexity. Supporting the actual payment landscape of Asia is what drives the node count up.</p>
<p>An engineer who understands this domain looks at a 60-node dependency graph and thinks “honestly, that seems about right.” An engineer who doesn’t understand the domain either proposes “simplifications” that would break the business, or sits paralysed by complexity they can’t reason about because they don’t understand why each node exists.</p>
<p>As the basketball coach John Wooden observed, “It’s what you learn after you know it all that counts.” The engineer who <em>thinks</em> they understand the booking flow because they’ve read the code is the one most likely to simplify away something load-bearing. The engineer who understands the business knows which complexity is essential and which is debt.</p>
<p>This principle applies far beyond travel. Every domain has hidden complexity that looks like over-engineering from the outside:</p>
<p>Healthcare has billing codes, insurance networks, HIPAA compliance, and prior authorisation workflows that make a simple “book an appointment” flow look like a state machine from a graduate textbook. E-commerce has inventory management across warehouses, return logistics, and promotional stacking rules where the interaction effects between discounts can surprise even the people who designed them. Insurance has actuarial models, claims adjudication processes, and regulatory capital requirements that vary by jurisdiction and product type.</p>
<p>The complexity that kills you isn’t in the code — it’s in the rules the code encodes. And if the people writing the code don’t understand the rules, they’ll encode them wrong.</p>
<h2>When “Booking” Means Five Different Things</h2>
<p>Eric Evans’ concept of Ubiquitous Language from Domain-Driven Design sounds academic until you’ve lived through what happens without it.</p>
<p>At an OTA, “booking” could mean: the act of a customer clicking “confirm” (the user event), the record in the booking database (the data entity), the financial obligation created (the accounting event), the reservation held at the hotel (the partner-side commitment), or the entire lifecycle from search to stay to review (the product journey). When the business says “we need to fix the booking flow,” which one do they mean? When an engineer builds a “booking service,” which definition are they encoding?</p>
<p>This isn’t pedantic. Ambiguity in language becomes ambiguity in architecture. If two teams use “booking” to mean different things, their services will model different concepts — and the integration between them will be where bugs live.</p>
<p>One area where Agoda has genuinely excelled — and has for a long time — is maintaining a ubiquitous language that runs from business meetings to database tables. If a business person says “rate plan” or “rate channel,” you can trace those exact terms from the frontend UI through the service layer all the way down to the database. The business vocabulary and the technical vocabulary are the same vocabulary. This isn’t a recent initiative or a DDD project someone championed — it’s been a long-standing cultural norm. And that consistency is one of the reasons our engineers <em>can</em> reason about the business, because the code literally speaks the same language as the business.</p>
<p>We use DDD extensively for our service architecture — a good part of the planning for our monolith split was guided by DDD principles. But the hardest part isn’t the framework. It’s getting engineers to internalise the <em>skill</em> of domain identification rather than memorising the current domain map. The common failure mode is engineers wanting to be <em>told</em> what the domains are so they can move on, rather than learning how to identify domains themselves. Knowing the current boundaries isn’t enough. Engineers need to understand how to spot when a new domain is emerging and how to classify new features into the right place. Without that skill, the domain model ossifies — it was right when it was drawn, but the business has moved on and nobody noticed because nobody knew how to re-evaluate.</p>
<p>Sound familiar? It’s the same pattern as the PO story. Engineers want the answer (“tell me the domains”) rather than the understanding (“teach me how to see domains”). The convenient shortcut undermines the deeper capability.</p>
<p>A deeper dive into how we apply Domain-Driven Design to service decomposition is coming in a future post.</p>
<h2>Domain Fluency Changes Engineering Decisions (Not Just Business Ones)</h2>
<p>Here’s a story that reframes what domain understanding actually looks like in practice.</p>
<p>A team needed to build complex dynamic pricing and commission logic — the kind of business rules that live at the intersection of pricing strategy, partner relationships, and search algorithms. Based on multiple variables, the system would adjust commission and price, which would then flow through to affect ranking. Pure domain complexity.</p>
<p>The engineers who built it had deep domain understanding. And that understanding manifested in an unexpected way: they chose Cucumber — BDD-style testing with a domain-specific language — to describe the behaviour.</p>
<p>Now, I’m generally not a fan of Cucumber. I think it’s great in theory and fails in practice for most teams. But this time it worked beautifully. The engineers wrote complex Cucumber tests using the DSL to precisely describe the pricing and commission logic in business-readable language. This wasn’t just a test suite — it became a communication artifact. The team that owned the downstream system could read it and verify the logic. The PO could read it and confirm the business rules. Everyone was on the same page, expressed in a shared language that was both executable and human-readable.</p>
<p>An engineer who doesn’t understand the domain would have written unit tests that verified the code worked. These engineers understood the domain well enough to recognise that the real risk wasn’t “does the code work?” but “does everyone agree on what the code should do?” — and they chose a tool specifically designed to solve that communication problem.</p>
<p>Domain-fluent engineers don’t just build better features — they make better <em>engineering</em> decisions. Tooling choices, testing strategies, API design, service boundaries, documentation approaches. Domain fluency isn’t a separate skill from technical skill. It’s what makes technical skill commercially relevant.</p>
<p>That Cucumber DSL was, in effect, the ubiquitous language for the pricing domain. It worked because the engineers understood the domain well enough to express it in terms both code and humans could parse. This is exactly what Evans advocates — and it happened not because someone mandated DDD, but because domain-fluent engineers naturally gravitated toward the right tool.</p>
<h2>The Cognitive Ceiling (Or: Why You Can’t Know Everything)</h2>
<p>Before you conclude that every engineer should understand the entire business, let me complicate the picture.</p>
<p>In the early days at Agoda, I led the NHA project — Non-Hotel Accommodation, essentially competing with Airbnb. The project required touching <em>every</em> area of the accommodation side of the business: pricing, booking, content — plus building entirely new domains on top. The team had to hold multiple complex domains in their heads simultaneously.</p>
<p>The project was hard, and the team moved slowly at times. Not because they were bad — because you simply can’t hold pricing logic, booking flows, content management, <em>and</em> new accommodation type rules in your head simultaneously without cognitive overload slowing you down. There’s a real upper bound on how much domain complexity one team can absorb.</p>
<blockquote>
<p>“The greatest enemy of knowledge is not ignorance, it is the illusion of knowledge.”* — Daniel J. Boorstin, historian and former Librarian of Congress*</p>
</blockquote>
<p>The surprising upside: I had remarkably low turnover on that team in the early days. They were always working on something new. The cognitive load was high, but the variety kept engineers engaged. There’s a genuine tension between “specialise to manage complexity” and “variety keeps people motivated” that’s worth naming honestly.</p>
<p>This is one of the underappreciated arguments for domain-driven service boundaries. It’s not just about independent deployment or technical decoupling. It’s about keeping the domain scope per team small enough that engineers can genuinely understand it. Domain separation isn’t just a technical architecture decision — it’s a <em>cognitive</em> architecture decision. You split systems along domain lines partly so that no single team has to hold too many business contexts simultaneously.</p>
<p>The answer isn’t “every engineer should understand the whole business.” The answer is “every engineer should deeply understand <em>their</em> domain, and the organisation should structure teams so that’s achievable.” Rotation through domains builds the cross-domain pattern recognition that’s invaluable, but trying to hold every domain simultaneously just makes you slow and stressed.</p>
<h2>Building Domain Fluency: What Actually Works</h2>
<h2>For Individual Engineers</h2>
<p><strong>Shadow the business.</strong> Sit in on a sales call. Watch a support agent handle a complaint. Attend a partner onboarding session. One afternoon of shadowing teaches you more about the domain than a month of reading documentation. At Agoda, we have a programme that lets engineers shadow call centre agents for a day — hearing real customers struggle with real problems in real time. Our UX researchers conduct interviews with partners and publish the recordings internally so engineers can watch them on their own schedule. These aren’t senior engineer perks. They’re available from day one.</p>
<p><strong>Follow the money.</strong> For any product company: How does a dollar enter the system? How does it move between entities? Where does it exit? Following the money reveals the real architecture of the business — the one that matters more than your service dependency diagram.</p>
<p><strong>Ask “why does this rule exist?”</strong> every single time you encounter a business rule that seems arbitrary. The answer is always one of: regulatory requirement, partner contract, historical accident, or “nobody knows.” Each answer tells you something important, and “nobody knows” is the most important answer of all — because it means that rule is either load-bearing for a reason nobody documented or it’s dead weight nobody’s had the courage to remove.</p>
<p><strong>Learn the vocabulary before the codebase.</strong> When joining a new team, spend your first week learning the domain language, not only the code. The code will make more sense once you understand what it’s trying to model. If your company has a ubiquitous language — terms that flow from business meetings to database columns — learn that language first. It’s the Rosetta Stone for your codebase.</p>
<p><strong>Read the financial reports.</strong> If your company is public, the annual report tells you how the company makes money, what the risks are, and what the strategic priorities are. This is the business context that never makes it into Jira tickets but shapes every decision about what gets built.</p>
<h2>For Engineering Leaders</h2>
<p><strong>Invest in domain onboarding.</strong> Your new hire bootcamp covers the tech stack, the CI/CD pipeline, and the coding standards. Add the business model, the key business processes, and the regulatory environment. If you’re in travel, explain how OTA pricing actually works. If you’re in fintech, walk through the payment lifecycle. An engineer who understands the business in week one makes better decisions in month one.</p>
<p><strong>Create domain days.</strong> Regular sessions where business experts explain a part of the domain to engineering. Not a presentation — a Q&amp;A. Let engineers ask the “obvious” questions that reveal critical assumptions everyone else has internalised and forgotten they know.</p>
<p><strong>Don’t gate domain understanding behind seniority.</strong> If your juniors aren’t talking to users, they’re building a habit of <em>not</em> understanding the domain that becomes harder to break with every passing year. The engineer who’s been executing tickets for five years and then gets promoted to senior doesn’t magically develop domain fluency. They’ve spent five years practising the opposite. Include user interaction in junior engineer expectations from their first sprint.</p>
<p><strong>Rotate engineers across domains — deliberately.</strong> An engineer who’s only ever worked on the payment service has a narrow view of the business. Rotation builds the cross-domain pattern recognition that prevents integration failures and builds empathy between teams. But be deliberate about the cognitive load: a team that touches every domain simultaneously moves slowly. A team that rotates through domains sequentially builds breadth without overload.</p>
<p><strong>Include domain understanding in the career ladder.</strong> If your engineering levels only reward technical skill, you’re signalling that domain knowledge doesn’t matter. Senior and Staff engineers should be expected to demonstrate deep domain understanding — not as a nice-to-have, but as a core competency. At Agoda, we don’t treat engineers as ticket executors. The expectation is that you understand the domain you’re building for. That’s not a cultural bonus — it’s what lets us operate a platform with 300+ payment methods and suppliers ranging from global hotel chains to individual homes without everything collapsing.</p>
<p><strong>Make the business model visible.</strong> Create architecture diagrams that show the business flow, not just the technical flow. When engineers can see how their service fits into the value chain, they stop optimising in isolation and start making decisions that serve the whole system.</p>
<h2>The Bottom Line</h2>
<p>As Peter Drucker observed, “There is nothing so useless as doing efficiently that which should not be done at all.” We’ve built an industry around making engineers more technically proficient — better at building things fast, building things that scale, building things that don’t break. And we’ve systematically underinvested in helping them understand what to build and why.</p>
<p>The code compiles. The tests pass. The feature ships. And the business loses money because the engineer who built the commission logic didn’t understand why that 2% difference matters in the Japanese market, or the engineer who “simplified” the booking flow didn’t realise that the node they removed was handling a regulatory requirement for a specific payment method in a specific jurisdiction.</p>
<p>Domain understanding isn’t the opposite of technical excellence. It’s the thing that makes technical excellence worth having. An engineer who understands the business writes code that models reality. An engineer who doesn’t writes code that models their best guess about reality — and best guesses, compounded across hundreds of engineers and thousands of commits, are how you end up with systems that technically work and commercially fail.</p>
<p>The way knowledge flows to engineers matters more than the volume. Too little, and they build blind. Too much in the wrong form, and they never develop sight. The goal isn’t a PO on tap or a wiki nobody reads. It’s engineers who understand the domain deeply enough to reason independently — who can look at a 60-node dependency graph of the booking systems and know which nodes are essential and which are debt, who can choose Cucumber over unit tests because they understand the communication problem matters more than the coverage metric, who can push back on a requirement not because they’re difficult but because they know the domain well enough to spot the contradiction.</p>
<p>You don’t have a technical problem. You have a comprehension problem. And the good news is that comprehension problems, unlike some technical problems, are solvable with intention, structure, and the willingness to let engineers get close enough to the business to understand it.</p>
<p>Now, if you’ll excuse me, I have a Definition of Ready to go dismantle. Turns out well-formatted tickets aren’t a substitute for engineers who understand what they’re building — but try explaining that to a team that’s been treating Jira like a food menu.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Stop Copying Prod Into Dev: Test Data Strategies That Actually Scale</title>
    <link>https://blog.dicko.dev/posts/stop-copying-prod-into-dev-test-data-strategies-that-actually-scale/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/stop-copying-prod-into-dev-test-data-strategies-that-actually-scale/</guid>
    <pubDate>Tue, 03 Mar 2026 11:40:33 GMT</pubDate>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>software-testing</category>
    <category>microservices</category>
    <category>automation-testing</category>
    <description>Or: Your Database Snapshots Are a Confession, Not a Solution</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/stop-copying-prod-into-dev-test-data-strategies-that-actually-scale/cover.webp" alt=""></figure>
<p>It was April in Phaya Thai, and the aircon had died again. Three pedestal fans oscillated uselessly across the open-plan floor, pushing warm air from one side of the room to the other like a bureaucracy redistributing blame. I was staring at a progress bar — a database restore from production, creeping toward completion at the pace of a man who knows he’s early for a meeting he doesn’t want to attend. Forty minutes in, maybe twenty to go. The AC technician had been called. The DB restore had been started. Neither was going to finish fast enough to matter. Somewhere behind me, someone racked up the pool table, because what else do you do when your entire development environment is held hostage by a backup file that’s grown three gigabytes since anyone last checked?</p>
<p>That restore was the whole workflow. Not a workaround, not a temporary measure — it was how we got data into dev. Copy prod, strip out the bits that looked obviously sensitive, cross your fingers, and wait. And when the question inevitably surfaced years later in a different office, a different country, a different Slack channel — “Is there any way we can periodically take prod data snapshots and use them to seed our test environments?” — I recognised it instantly. Not as a bad question. As a confession. The same confession I’d been making back in that sweltering Phaya Thai office: we’ve lost the ability to describe our own data in code. And once you’ve lost that, everything downstream — your tests, your onboarding, your confidence in deployments — starts to rot from the inside.</p>
<p>As the physicist Richard Feynman once observed, “The first principle is that you must not fool yourself — and you are the easiest person to fool.” We’ve been fooling ourselves into thinking that production snapshots are a reasonable default. They’re not. They’re a crutch. And sometimes you genuinely need a crutch — but if you’re building a new system and reaching for one on day one, something’s already broken.</p>
<p>This post is more technical than my usual fare. There’s code. Quite a lot of it, actually. We’ll walk through working examples in both .NET and Spring Boot, covering how to eliminate lookup tables, seed data programmatically, build fluent test data factories, and run real databases in tests without shared state. If you’ve ever spent twenty minutes waiting for a Docker container stuffed with a QA database dump to finish starting up so you could run your tests, this one’s for you.</p>
<h2>Why Prod Snapshots Hurt More Than They Help</h2>
<p>Let’s enumerate the damage before we talk about alternatives, because the costs are easy to wave away when the snapshot is already sitting right there in your pipeline:</p>
<p><strong>PII risk.</strong> Prod databases contain real customer data. Copying it to dev environments — even “anonymised” — creates compliance surface area that somebody has to manage. Anonymisation is harder than it sounds, and “good enough” anonymisation tends to decay over time as schemas change and new PII fields slip through the cracks.</p>
<p><strong>Scale mismatch.</strong> Prod databases are enormous. Dev environments are not. You either copy a subset (which might miss the exact data shape you need) or the full thing (which takes hours and costs money nobody budgeted for).</p>
<p><strong>Stale data.</strong> Snapshot from Tuesday, bug found on Friday. Is the snapshot still valid? Nobody knows. Nobody checks. Everybody assumes.</p>
<p><strong>Shared-state fragility.</strong> This one deserves a story.</p>
<p>We had teams at Agoda creating test hotels in the shared QA database, writing assertions against those specific hotels, then watching their tests start failing across teams because someone else was exploring the same QA environment and changed something about that hotel. The tests weren’t testing code — they were testing a specific data state that anyone could mutate at any time. It’s the software equivalent of writing your exam answers on a shared whiteboard and being surprised when someone erases them.</p>
<p>We also had systems using database restores from QA — Docker containers with the data baked in. They were growing to multiple gigabytes, taking minutes just to start up. When we moved to SQL scripts with Testcontainers, the entire test suite started <em>feeling like unit tests</em> — that’s how fast it got.</p>
<p>We measured the impact concretely. We tracked how many test suites engineers actually ran locally on their dev machines versus what we saw running in CI, grouped at the MR level. The gap represented tests that were too slow or too painful to run locally — the tests engineers skipped because they didn’t want to wait.</p>
<p>The starting point was Docker Compose with data restored from QA baked into the containers. Running tests meant dropping to a CLI outside the IDE, waiting for multi-gigabyte containers to spin up, and hoping nobody had mutated the QA data since the last image build. Around 10–20% of MRs showed local test runs. The rest went straight to CI and waited.</p>
<p>We moved two things in one step: replaced the QA data restores with SQL scripts, and switched to Testcontainers. That combination made the integration tests fast and self-contained enough that we could add them to the default “Run All Tests” action in the IDE — something that wasn’t feasible with the old Docker Compose setup. Within a few weeks, local test execution jumped to 70% of MRs on some days running integration tests locally. Not because we told engineers to run their tests. Because we made it easy enough that they stopped skipping them.</p>
<figure><img src="https://blog.dicko.dev/posts/stop-copying-prod-into-dev-test-data-strategies-that-actually-scale/image_2.webp" alt="" loading="lazy"></figure>
<h2>The Principle: Your API Owns Its Data</h2>
<p>Before we dive into techniques, let’s establish the principle that makes all of this work: <strong>if your API owns its data, it should be able to describe that data completely in code.</strong></p>
<p>This connects directly to Domain-Driven Design and the bounded context concept — and if you think that’s ivory-tower talk, <a href="https://gist.github.com/chitchcock/1281611">Bezos issued the same mandate at Amazon</a> back in 2002. Your service’s database is an implementation detail of your API. If another team needs your data, they get it through your API, not by reading your tables. And if your API can serve the data, your code already knows what valid data looks like.</p>
<p>This means reference data — statuses, types, categories — should be defined in code, not manually inserted into lookup tables. Schema should be version-controlled through migrations, not applied by hand. Seed data for dev and test should be generated by the same code that validates it in production.</p>
<p>When this principle holds, you never need a prod snapshot for development. Your code <em>is</em> the authoritative description of your data.</p>
<p>As the management theorist W. Edwards Deming put it, “If you can’t describe what you are doing as a process, you don’t know what you’re doing.” If your API can’t describe its own data, what does that tell you?</p>
<h2>Technique 1: Kill the Lookup Table (Or At Least Own It)</h2>
<p>Lookup tables are one of the primary reasons teams reach for prod snapshots. “We need the StatusTypes table populated with the right IDs.” It’s the siren song of copying prod, and it starts with this one innocent dependency.</p>
<p>Here’s the progression of maturity:</p>
<h3>Level 0: Manual Lookup Tables (The Problem)</h3>
<p>Someone ran an INSERT script against prod years ago. Nobody’s sure if the dev database has the same values. The IDs might be different between environments. There’s no version control. This is where the snapshot reflex is born.</p>
<h3>Level 1: Enum-Backed Lookup Tables with Auto-Seeding</h3>
<p>Your code defines the enum. Your migrations create and populate the table. The database is always in sync with the code.</p>
<p><strong>EF Core (.NET):</strong></p>
<pre><code>// Define the enum as the source of truth
public enum BookingStatus
{
    Pending = 1,
    Confirmed = 2,
    Cancelled = 3,
    Completed = 4,
    Refunded = 5
}
// Create a lookup entity that mirrors the enum
public class BookingStatusLookup
{
    public int Id { get; set; }
    public string Name { get; set; } = string.Empty;
}
// In your DbContext OnModelCreating
protected override void OnModelCreating(ModelBuilder modelBuilder)
{
    // Seed the lookup table from the enum - always in sync
    modelBuilder.Entity&lt;BookingStatusLookup&gt;().HasData(
        Enum.GetValues&lt;BookingStatus&gt;()
            .Select(e =&gt; new BookingStatusLookup
            {
                Id = (int)e,
                Name = e.ToString()
            })
            .ToArray()
    );
    // Your booking entity uses the enum directly
    modelBuilder.Entity&lt;Booking&gt;()
        .Property(b =&gt; b.Status)
        .HasConversion&lt;int&gt;();  // Stored as int, used as enum
}
</code></pre>
<p><strong>Spring Boot / JPA:</strong></p>
<pre><code>public enum BookingStatus {
    PENDING, CONFIRMED, CANCELLED, COMPLETED, REFUNDED
}

@Entity
public class Booking {
    @Id @GeneratedValue
    private Long id;
    @Enumerated(EnumType.STRING)
    private BookingStatus status;
}
</code></pre>
<p>With Flyway or Liquibase migrations, the schema and constraints are always derived from your code and version-controlled. You’ll see spring.jpa.hibernate.ddl-auto=create in tutorials — it&#39;s fine for local prototyping, but in real services it produces &quot;works on my machine&quot; schemas that drift between environments and miss indices or constraints that migrations would express. For anything beyond a throwaway spike, let your migration tool own the schema.</p>
<blockquote>
<p><em><strong>A word on enum stability.</strong></em>* The moment you seed a database from an enum, the numeric values become data. Treat them like a public API. Always assign explicit values (**Pending = 1, not just **Pending). Never reorder enum members — if **Cancelled = 3 has been written to 50,000 rows, reordering it doesn&#39;t update those rows. Never delete enum values — mark them obsolete but keep the numeric value reserved. EF Core&#39;s **HasData will generate a DELETE migration if the value disappears, which will FK-violate against existing rows. In Spring/JPA, prefer **EnumType.STRING over *<em>EnumType.ORDINAL for exactly this reason — ordinal values break when enums are reordered, while string values are stable.</em></p>
</blockquote>
<h3>Level 2: Smart Enums — No Lookup Table At All</h3>
<p>If your API owns its data and no other service reads your database directly, you may not need a lookup table at all. The enum <em>is</em> the reference data.</p>
<p><strong>Ardalis SmartEnum (.NET):</strong></p>
<pre><code>public sealed class BookingStatus : SmartEnum&lt;BookingStatus&gt;
{
    public static readonly BookingStatus Pending = new(nameof(Pending), 1);
    public static readonly BookingStatus Confirmed = new(nameof(Confirmed), 2);
    public static readonly BookingStatus Cancelled = new(nameof(Cancelled), 3);
    public static readonly BookingStatus Completed = new(nameof(Completed), 4);

    // You can add behaviour directly on the enum
    public bool CanTransitionTo(BookingStatus target) =&gt;
        (this, target) switch
        {
            (_, _) when this == target =&gt; false,
            var (from, to) when from == Pending
                =&gt; to == Confirmed || to == Cancelled,
            var (from, to) when from == Confirmed
                =&gt; to == Completed || to == Cancelled,
            _ =&gt; false
        };
    private BookingStatus(string name, int value) : base(name, value) { }
}

// EF Core value conversion — stores as int, no lookup table needed
modelBuilder.Entity&lt;Booking&gt;()
    .Property(b =&gt; b.Status)
    .HasConversion(
        s =&gt; s.Value,
        v =&gt; BookingStatus.FromValue(v)
    );
</code></pre>
<p>The API returns &quot;Confirmed&quot; to the client. The database stores 2. The code defines both. No lookup table, no seeding, no prod snapshot needed.</p>
<h3>The DBA Pushback (And Why It’s Half Right)</h3>
<p>The line “lookup tables exist because databases predate APIs” is deliberately provocative. It needs to be, because the default assumption in most organisations is that lookup tables are obviously correct, and that assumption goes unexamined. But DBAs will push back, and some of their arguments are legitimate.</p>
<p><strong>“What about referential integrity?”</strong> This is the strongest argument — but only if the database has multiple writers. In a microservices architecture where your API is the only writer to its database, you already control all writes. The enum in your code <em>is</em> the constraint — enforced at compile time, which is strictly stronger than a runtime FK violation. You can’t even express an invalid value in a language with a proper type system.</p>
<p>If you want both, use Level 1: define the enum in code, auto-seed the lookup table from the enum via migrations, and let the FK exist as a belt-and-suspenders safety net. The key is that the enum is the source of truth, and the table is derived from it. Not the other way around.</p>
<p><strong>“What about reporting?”</strong> When an analyst runs SELECT status_id, COUNT(*) FROM bookings GROUP BY status_id, they see numbers. Fair. Three responses: store as string instead of int (EF Core&#39;s EnumToStringConverter and JPA&#39;s @Enumerated(EnumType.STRING) handle this trivially); auto-seed the lookup from the enum as a view rather than independent data; or — the microservices answer — reporting should hit a read model, not your service database. If reporting queries hit your transactional database directly, you have the &quot;Reach-in Reporting&quot; antipattern that Mark Richards describes in <em>Microservices AntiPatterns and Pitfalls</em>.</p>
<p><strong>“So when do lookup tables win?”</strong> When the values are user-editable at runtime — product categories a business user can add without a code deploy. When the values carry metadata beyond name and ID — a Currency table with symbol, decimal places, active flag. When multiple applications write to the same database. When your reporting layer needs to join against them in a read model.</p>
<p>The argument isn’t “never use lookup tables.” It’s “stop treating them as the default, and start asking who consumes this database.”</p>
<h2>Technique 2: Programmatic Data Seeding</h2>
<h3>The Maturity Ladder</h3>
<p>Level Practice Reproducible? PII Risk 0 Shared prod clone No — point in time High 1 SQL scripts in repo Partially — may not be idempotent None 2 Migration-based seeds (HasData, Flyway) Yes — deterministic None 3 Code-driven seeding + fluent builders Yes — type-safe None 4 Ephemeral DB per test suite (Testcontainers) Yes — clean slate None</p>
<p>Most teams are at Level 0 or 1 and think they need to jump to Level 4. They don’t — each level is an improvement. Moving from Level 0 to Level 2 alone eliminates PII risk and makes environments reproducible.</p>
<p><strong>EF Core HasData (for migration-managed seed data):</strong></p>
<pre><code>// In OnModelCreating — generates INSERT statements in migrations
modelBuilder.Entity&lt;Currency&gt;().HasData(
    new Currency { Id = 1, Code = &quot;USD&quot;, Name = &quot;US Dollar&quot;, Symbol = &quot;$&quot; },
    new Currency { Id = 2, Code = &quot;THB&quot;, Name = &quot;Thai Baht&quot;, Symbol = &quot;฿&quot; },
    new Currency { Id = 3, Code = &quot;EUR&quot;, Name = &quot;Euro&quot;, Symbol = &quot;€&quot; }
);
</code></pre>
<p>HasData requires explicit primary keys and is designed for data that changes rarely. It generates migration scripts, so it&#39;s version-controlled and reproducible.</p>
<p><strong>IHostedService seeding (EF Core 6/7/8 — the pattern most teams actually use):</strong></p>
<pre><code>public class DatabaseSeeder : IHostedService
{
    private readonly IServiceProvider _serviceProvider;
    public DatabaseSeeder(IServiceProvider serviceProvider)
        =&gt; _serviceProvider = serviceProvider;
    public async Task StartAsync(CancellationToken cancellationToken)
    {
        using var scope = _serviceProvider.CreateScope();
        var context = scope.ServiceProvider
            .GetRequiredService&lt;BookingDbContext&gt;();
        await context.Database.MigrateAsync(cancellationToken);
        if (!await context.Currencies.AnyAsync(cancellationToken))
        {
            context.Currencies.AddRange(
                new Currency { Code = &quot;USD&quot;, Name = &quot;US Dollar&quot;, Symbol = &quot;$&quot; },
                new Currency { Code = &quot;THB&quot;, Name = &quot;Thai Baht&quot;, Symbol = &quot;฿&quot; },
                new Currency { Code = &quot;EUR&quot;, Name = &quot;Euro&quot;, Symbol = &quot;€&quot; }
            );
            await context.SaveChangesAsync(cancellationToken);
        }
    }
    public Task StopAsync(CancellationToken ct) =&gt; Task.CompletedTask;
}
// In Program.cs
builder.Services.AddHostedService&lt;DatabaseSeeder&gt;();
</code></pre>
<p>This runs on every application startup, applies pending migrations, and seeds data idempotently. It’s the pragmatic choice for teams not yet on EF Core 9.</p>
<blockquote>
<p><em><strong>A word of caution:</strong> <strong>register this seeder in development and test environments only. If someone cargo-cults it into production without an environment guard, you’ve introduced implicit startup-time database mutation — and if two pods start simultaneously, you’ve got a race condition. In production, prefer migration tooling that runs outside the application process. Your <em><a href="https://blog.dicko.dev/posts/the-impact-of-paved-paths-and-embracing-the-future-of-development/"><em>paved path</em></a></em> templates should make the safe thing the default: seeder registered in</strong> Development, migrations applied via a pipeline step.</em></p>
</blockquote>
<p><strong>Spring Boot — Flyway migrations with seed data:</strong></p>
<pre><code>-- V1__create_currency_table.sql
CREATE TABLE currency (
    id SERIAL PRIMARY KEY,
    code VARCHAR(3) NOT NULL UNIQUE,
    name VARCHAR(50) NOT NULL,
    symbol VARCHAR(5) NOT NULL
);
-- V2__seed_currencies.sql
INSERT INTO currency (code, name, symbol) VALUES
    (&#39;USD&#39;, &#39;US Dollar&#39;, &#39;$&#39;),
    (&#39;THB&#39;, &#39;Thai Baht&#39;, &#39;฿&#39;),
    (&#39;EUR&#39;, &#39;Euro&#39;, &#39;€&#39;);
</code></pre>
<p>Flyway tracks which migrations have run. Add a new currency? Add a new migration. Version-controlled, reproducible, works identically in every environment.</p>
<h2>Technique 3: Fluent Test Data Builders</h2>
<p>This is where we get to the heart of the prod-snapshot reflex. The reason engineers reach for snapshots in tests is because constructing valid domain objects is painful. A Booking might need a Property, a Guest, a RatePlan, a PaymentMethod, and three Currency records before it’s valid. When the alternative is a single docker-compose up that gives you everything, of course people take the shortcut.</p>
<p>The answer: builder patterns that make constructing complex test data trivial.</p>
<p><strong>The .NET builder pattern:</strong></p>
<pre><code>public class BookingBuilder
{
    private Guid _id = new(&quot;00000000-0000-0000-0000-000000000001&quot;);
    private string _guestName = &quot;Test Guest&quot;;
    private DateTime _checkIn = new(2025, 6, 1);
    private BookingStatus _status = BookingStatus.Pending;
    private string _currency = &quot;USD&quot;;
    private int _nights = 3;
    private decimal _amount = 450.00m;
    private int _propertyId = 1;

    public BookingBuilder WithId(Guid id)
        { _id = id; return this; }
    public BookingBuilder WithGuest(string name)
        { _guestName = name; return this; }
    public BookingBuilder CheckingInOn(DateTime date)
        { _checkIn = date; return this; }
    public BookingBuilder WithStatus(BookingStatus status)
        { _status = status; return this; }
    public BookingBuilder InCurrency(string currency)
        { _currency = currency; return this; }
    public BookingBuilder ForNights(int nights)
        { _nights = nights; return this; }
    public BookingBuilder WithAmount(decimal amount)
        { _amount = amount; return this; }
    public BookingBuilder AtProperty(int propertyId)
        { _propertyId = propertyId; return this; }

    public Booking Build() =&gt; new()
    {
        Id = _id,
        GuestName = _guestName,
        CheckIn = _checkIn,
        CheckOut = _checkIn.AddDays(_nights),
        Status = _status,
        TotalAmount = _amount,
        Currency = _currency,
        PropertyId = _propertyId
    };

    public static BookingBuilder ABooking() =&gt; new();
}

// Usage in tests — reads like English, every value is predictable
var booking = BookingBuilder.ABooking()
    .WithStatus(BookingStatus.Confirmed)
    .InCurrency(&quot;THB&quot;)
    .ForNights(7)
    .WithAmount(2800.00m)
    .Build();

// Need a cancelled booking? Override only what matters
var cancelled = BookingBuilder.ABooking()
    .WithStatus(BookingStatus.Cancelled)
    .Build();
// Everything else — guest name, check-in date, amount —
// uses sensible defaults. Deterministic. Every time.
</code></pre>
<p>Notice: no Faker, no Random, no Guid.NewGuid(). Every default is an explicit, known value. When a test creates a booking without specifying an amount, it gets 450.00 — not a random number between 50 and 5,000 that might trigger a different code path on Tuesday than it did on Monday.</p>
<p>This is a deliberate choice. Random test data is the enemy of deterministic tests. If your test passes with a random amount of 150.00 but fails with 5,001.00 because it hits a threshold you forgot about, you’ve got a flaky test that will erode trust in your entire suite. Tests should be boring. Tests should be predictable. The only values that vary should be the ones the test is explicitly exercising.</p>
<p><strong>Spring Boot / Java equivalent:</strong></p>
<pre><code>public class BookingBuilder {
    private String guestName = &quot;Test Guest&quot;;
    private LocalDate checkIn = LocalDate.of(2025, 6, 1);
    private int nights = 3;
    private BookingStatus status = BookingStatus.CONFIRMED;
    private BigDecimal totalAmount = new BigDecimal(&quot;450.00&quot;);
    private String currency = &quot;THB&quot;;

    public static BookingBuilder aBooking() {
        return new BookingBuilder();
    }

    public BookingBuilder withGuest(String name)
        { this.guestName = name; return this; }
    public BookingBuilder checkingInOn(LocalDate date)
        { this.checkIn = date; return this; }
    public BookingBuilder withStatus(BookingStatus status)
        { this.status = status; return this; }
    public BookingBuilder withAmount(BigDecimal amount)
        { this.totalAmount = amount; return this; }
    public BookingBuilder inCurrency(String currency)
        { this.currency = currency; return this; }
    public BookingBuilder forNights(int nights)
        { this.nights = nights; return this; }

    public Booking build() {
        return Booking.builder()
            .guestName(guestName)
            .checkIn(checkIn)
            .checkOut(checkIn.plusDays(nights))
            .status(status)
            .totalAmount(totalAmount)
            .currency(currency)
            .build();
    }
}

// Same pattern, same determinism
var booking = BookingBuilder.aBooking()
    .withStatus(BookingStatus.PENDING)
    .inCurrency(&quot;USD&quot;)
    .forNights(5)
    .withAmount(new BigDecimal(&quot;750.00&quot;))
    .build();
</code></pre>
<p>The key message here is uncomfortable but important: <strong>if creating valid test data is harder than copying prod, the problem is your test data tooling, not the absence of prod data.</strong></p>
<h2>Why Not Bogus/Faker?</h2>
<p>Libraries like Bogus and Java Faker are popular, and you’ll see them recommended everywhere. They generate realistic-looking random data — names, addresses, amounts, dates. They’re fun to use. They also introduce non-determinism into your tests, which is another way of saying they introduce flakiness that you can’t reproduce.</p>
<p>A test that fails in CI but passes locally because Faker generated a different value is worse than a test that doesn&#39;t exist. At least the missing test isn&#39;t lying to you. Random data means your tests exercise different code paths on different runs. That&#39;s not testing — that&#39;s hoping.</p>
<p>If you need to test multiple data shapes, write multiple test cases with explicit values. It’s more code, yes. It’s also code that fails the same way every time, which means you can actually debug it.</p>
<h2>Technique 4: Testcontainers — Real Databases, Zero Shared State</h2>
<p>This is where everything comes together. Testcontainers gives you ephemeral, real database instances per test suite — your tests run against the same database engine as production, with data you control completely.</p>
<p><strong>The .NET pattern (WebApplicationFactory + Testcontainers):</strong></p>
<pre><code>public class BookingApiFactory
    : WebApplicationFactory&lt;Program&gt;, IAsyncLifetime
{
    private readonly PostgreSqlContainer _postgres =
        new PostgreSqlBuilder()
            .WithImage(&quot;postgres:16-alpine&quot;)
            .Build();

    protected override void ConfigureWebHost(IWebHostBuilder builder)
    {
        builder.ConfigureTestServices(services =&gt;
        {
            // Remove the real database registration
            var descriptor = services.SingleOrDefault(
                d =&gt; d.ServiceType ==
                    typeof(DbContextOptions&lt;BookingDbContext&gt;));
            if (descriptor != null) services.Remove(descriptor);

            // Replace with Testcontainers connection
            services.AddDbContext&lt;BookingDbContext&gt;(options =&gt;
                options.UseNpgsql(
                    _postgres.GetConnectionString()));
        });
    }

    public async Task InitializeAsync()
    {
        await _postgres.StartAsync();

        // Apply migrations — your schema is always current
        using var scope = Services.CreateScope();
        var context = scope.ServiceProvider
            .GetRequiredService&lt;BookingDbContext&gt;();
        await context.Database.MigrateAsync();
    }

    public new async Task DisposeAsync()
        =&gt; await _postgres.DisposeAsync();
}
</code></pre>
<p><strong>The test class — clean, focused, no shared state:</strong></p>
<pre><code>public class BookingApiTests
    : IClassFixture&lt;BookingApiFactory&gt;
{
    private readonly HttpClient _client;
    private readonly BookingDbContext _db;

    public BookingApiTests(BookingApiFactory factory)
    {
        _client = factory.CreateClient();
        var scope = factory.Services.CreateScope();
        _db = scope.ServiceProvider
            .GetRequiredService&lt;BookingDbContext&gt;();
    }

    [Fact]
    public async Task CreateBooking_ReturnsCreated()
    {
        // Arrange — use your builders, not prod data
        var request = BookingBuilder.ABooking()
            .WithStatus(BookingStatus.Pending)
            .InCurrency(&quot;USD&quot;)
            .Build();

        // Act
        var response = await _client
            .PostAsJsonAsync(&quot;/api/bookings&quot;, request);

        // Assert
        response.StatusCode.Should().Be(HttpStatusCode.Created);
        var created = await response.Content
            .ReadFromJsonAsync&lt;Booking&gt;();
        created.Status.Should().Be(BookingStatus.Pending);
    }
}
</code></pre>
<p>For Spring Boot, the pattern is nearly identical: @Testcontainers annotation, @Container for the PostgreSQL container, @DynamicPropertySource to wire the connection string, and @SpringBootTest with RANDOM_PORT. The key insight is the same: your test owns its database, your migrations run on startup, your builders provide the data.</p>
<p><strong>Why this beats a prod snapshot or shared QA database:</strong></p>
<p>Each test suite gets a fresh database — no “someone changed my test hotel” problems. Migrations are tested every time — schema drift is caught immediately. Tests run in CI with no external dependencies beyond Docker. Data is deterministic. No PII in dev environments.</p>
<p>We run Testcontainers in Docker-in-Docker in our CI; our CI network is 20Gbps so the image pull is fast, though the Docker layer cache not caching is a known annoyance we’re actively working on. Real-world friction exists. But the alternative — maintaining multi-gigabyte snapshot containers — creates more friction, not less.</p>
<p>Two gotchas worth mentioning before you hit them in production CI: first, parallel test execution can spin up too many containers and exhaust ports or memory — start one container per test suite, not per test. Second, if your CI runners don’t support privileged Docker, Testcontainers will fail entirely. In that case, fall back to a shared ephemeral database per pipeline stage, provisioned and torn down by the pipeline itself. It’s not as clean as per-suite isolation, but it’s still miles ahead of a shared QA database that anyone can mutate.</p>
<h2>When a Snapshot Is Actually the Right Tool</h2>
<p>Before we get to the decision framework, let’s be clear: this post is not anti-snapshot. Snapshots are a tool, and like any tool, the problem isn’t the tool — it’s reaching for it by default instead of by design.</p>
<h2>The Decision Framework</h2>
<p>When you’re staring at a new test file and wondering how to get data into it, here’s the decision tree:</p>
<p><strong>Is this a new system you’re building?</strong> Seed from code. Period. Define your reference data in enums, use HasData or Flyway for seed data, build test data factories, run Testcontainers for integration tests. You have no excuse and no legacy to blame.</p>
<p><strong>Is this an existing system with code-first schema management?</strong> You should still seed programmatically. If your seeding is missing data, fix the seeding — that’s a bug, not a reason to clone prod.</p>
<p><strong>Is this a legacy system where you don’t fully control the schema?</strong> Years ago, I was at a company where we took over systems from another outfit. 248 tables, and we only had more features to add. Getting the schema into source control was the first barrier — at least we could control versioning. But getting all the data we needed for test cases in the right structure with the right relations? That’s what we’d lost control of. The quickest solution: take a backup, anonymise the PII, and there’s your test database. It only works at small scale, too — you can’t be doing that with a terabyte database.</p>
<p><strong>Do you need realistic data distributions for performance testing?</strong> Prod snapshot — anonymised — is the right tool. Synthetic data can’t replicate production data distributions for query optimisation. We still use snapshots for testing index changes at Agoda, and that’s the right call.</p>
<p><strong>Do you need to reproduce a specific production bug?</strong> Targeted snapshot of the relevant data. Reproduce. Fix. Delete the snapshot.</p>
<p>Even in the cases where snapshots are justified, use guardrails: anonymise PII before it leaves production, time-box the snapshot so it doesn’t become the permanent dev database, and document <em>why</em> you needed it. If the answer is “we can’t create valid test data,” that’s a problem to fix, not a state to accept.</p>
<h2>The Bottom Line</h2>
<p>The Slack message that started this post wasn’t a bad question. It was a perfectly rational question from someone dealing with real friction. The problem isn’t the question — it’s that our industry has normalised the answer.</p>
<p>The coach John Wooden once said, “If you don’t have time to do it right, when will you have time to do it over?” Every prod snapshot you take is a bet that you’ll never need to do this properly. And every month that passes, the bet gets more expensive. The snapshot grows. The PII surface area expands. The schema drifts further from the code. The tests get slower. The engineers stop running them.</p>
<p>The path out isn’t dramatic. It’s incremental. Move from Level 0 to Level 2 on the maturity ladder. Own your enums. Seed your reference data from code. Build one deterministic test data builder for your most common domain object. Set up one Testcontainers integration test. Measure how many tests your engineers actually run locally.</p>
<p>That 20% to 70% jump in local test execution we saw? It didn’t come from a heroic rewrite. It came from making the default workflow fast enough that people stopped skipping it. That’s the whole game. Make the right thing the easy thing, and the right thing happens.</p>
<p>Now, if you’ll excuse me, I need to go review a merge request that adds a new enum value to our booking status. The engineer helpfully put it in the middle of the list without an explicit integer value. At least our migration tests will catch it — because we actually <em>run</em> those now.</p>
]]></content:encoded>
  </item>
  <item>
    <title>The History Lesson Nobody Wanted: Why Your Team Is Living in 1985</title>
    <link>https://blog.dicko.dev/posts/the-history-lesson-nobody-wanted-why-your-team-is-living-in-1985/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/the-history-lesson-nobody-wanted-why-your-team-is-living-in-1985/</guid>
    <pubDate>Mon, 02 Mar 2026 08:40:55 GMT</pubDate>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>agile</category>
    <category>scrum</category>
    <category>engineering-management</category>
    <description>Or: How We Forgot What Agile Actually Meant and Built a Cargo Cult Instead</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/the-history-lesson-nobody-wanted-why-your-team-is-living-in-1985/cover.webp" alt=""></figure>
<p>She’d raised it in three consecutive retros. The same point, worded slightly differently each time, the way you rephrase a question when you suspect nobody heard it the first time. The deployment pipeline was brittle — two manual steps, a restart sequence that only worked if you did it before 4 PM Bangkok time, and a Slack message to a person who was usually on leave. The meeting room on level 8 was too warm, the way they always are when twelve people sit in a space designed for eight — that slightly stale mix of Coffee Mate from the kitchen down the hall, someone’s Thai iced tea perspiring onto the vinyl desk, and the faint sweetness of a half-opened bag of Lays that had been doing the rounds since someone came back from a trip to Japan. The scrum master wrote her point on a Post-It, stuck it in the “action items” column, and moved on to the next topic. The Post-It was still there six weeks later, curling at the edges, the adhesive giving up before the process did. She stopped raising it after that. Not loudly — she just stopped. The kind of silence that doesn’t register in any standup or velocity chart but changes the temperature of every meeting it sits in.</p>
<p>That silence is the real cost of ceremony without substance. Not the wasted hours, not the theatre — the moment an engineer decides the process isn’t listening and quietly stops talking. We’ve spent twenty-five years building an industry around the <em>ceremony</em> of Agile while systematically ignoring the <em>substance</em> of it. The manifesto’s own authors have been saying this for over a decade. Most of us just weren’t listening — we were too busy estimating in Fibonacci numbers.</p>
<h2>The Forty-Year Loop</h2>
<p>There is a peculiar amnesia in the software industry. Every fifteen years or so, a generation of engineers discovers that the way they’re building software doesn’t work — that heavy process produces heavy failures — and they rebel. They write manifestos. They form movements. They invent new names for old ideas. And then, slowly, the rebellion becomes the establishment, the establishment becomes the bureaucracy, and the next generation discovers that the way they’re building software doesn’t work.</p>
<p>As the philosopher George Santayana warned, “Those who cannot remember the past are condemned to repeat it.” He was talking about civilisations. He might as well have been talking about software methodology conferences.</p>
<p>The pattern goes like this: engineers find a pragmatic way of working. Management formalises it. The formalisation kills the thing that worked. Engineers rebel. Repeat.</p>
<p><strong>1970: The Misunderstanding.</strong> Winston Royce published “Managing the Development of Large Software Systems.” The paper is famous for its diagram of sequential phases — what we now call waterfall. Here’s the part that almost nobody mentions: Royce presented that diagram as an example of what <em>doesn’t work</em>. He wrote, on page two, that this approach “is risky and invites failure.” He spent the rest of the paper proposing iterative alternatives. The industry read page one and stopped.</p>
<p>The term “waterfall” doesn’t even appear in Royce’s paper. It was coined six years later by Bell and Thayer, referencing the very diagram Royce said was dangerous. We built an entire era of software development on a misreading.</p>
<p><strong>1985: The Mandate.</strong> The U.S. Department of Defense published DOD-STD-2167, mandating the waterfall model for military software contractors. The same approach Royce warned against became official government policy. The result was predictable: projects that took years, cost billions, and frequently failed. The Standish Group’s 1994 CHAOS Report found a 16% success rate for software projects — and defence projects were among the worst performers.</p>
<p>Note the mechanism: a pragmatic observation was misread, formalised, mandated, and scaled. The original insight was lost. The ceremony remained.</p>
<p><strong>2001: The Rebellion.</strong> By the late 1990s, practitioners were independently developing lightweight methodologies — Kent Beck’s Extreme Programming, Jeff Sutherland and Ken Schwaber’s Scrum, Alistair Cockburn’s Crystal. In February 2001, seventeen of them met at the Snowbird ski resort in Utah. They expected disagreement — as the original history page describes it, a bigger gathering of organisational anarchists would be hard to find. Instead, they found common ground. Over two days, they wrote four value statements and published them as the Manifesto for Agile Software Development.</p>
<p>The manifesto is 68 words long. It takes about thirty seconds to read. It contains no mention of sprints, story points, velocity, Scrum Masters, Product Owners, SAFe, or certifications. It is a statement of values, not a process.</p>
<p><strong>2005–Now: The Cargo Cult.</strong> And then the cycle repeated. Again.</p>
<p>The Scrum Alliance was founded. Certifications were created. A two-day course produced a “Certified ScrumMaster.” The Scaled Agile Framework packaged Agile into an enterprise-friendly structure with roles, ceremonies, and governance that looked — to anyone who’d lived through 1985 — remarkably like waterfall with better branding.</p>
<p>The manifesto’s authors watched it happen. Dave Thomas, one of the seventeen signatories, declared in 2014 that the word “agile” had been subverted to the point of meaninglessness and the agile community had become largely an arena for consultants and vendors. Martin Fowler warned in his 2018 Agile Australia keynote about “Faux Agile” — agile in name without practices or values. Robert “Uncle Bob” Martin observed that the agile movement had pushed so many project managers in that they’d pushed the programmers out. Ron Jeffries coined “Dark Scrum.” Ken Schwaber, co-creator of Scrum itself, left the Scrum Alliance in 2009 and later called SAFe “unSAFe at any speed.”</p>
<p>The mechanism was identical to 1985. A pragmatic insight was formalised, certified, mandated, and scaled. The original values were lost. The ceremony remained.</p>
<h2>The Bamboo Runway</h2>
<p>Richard Feynman described cargo cults in his 1974 Caltech commencement speech: Pacific islanders who’d seen military planes bring supplies during World War II built runways from bamboo and lit signal fires, mimicking the form of technology without understanding the substance. The planes never came.</p>
<p>You might be standing on a bamboo runway right now.</p>
<p>The manifesto says “Individuals and interactions over processes and tools.” Your organisation has a mandatory standup format, a mandatory retro format, and a mandatory planning poker tool. Deviation from the format is treated as a process failure.</p>
<p>The manifesto says “Working software over comprehensive documentation.” Your “Definition of Done” requires updating Confluence, Jira, and a release tracking spreadsheet. Nobody reads any of them.</p>
<p>The manifesto says “Customer collaboration over contract negotiation.” Requirements arrive as Jira tickets from a product committee. Engineers never talk to customers.</p>
<p>The manifesto says “Responding to change over following a plan.” Changing scope mid-sprint requires a “sprint change request.” Velocity is tracked as a performance metric.</p>
<p>Here’s the deeper test: look at who attends your Agile ceremonies. If the answer is “mostly non-engineers,” you’ve recreated 1985. The meetings exist for management visibility, not for the team building software.</p>
<p>And here’s an even simpler one: watch your engineers’ eyes during standup. Who do they look at when they talk? If it’s each other — trading context, flagging dependencies, offering help — you’ve got a team synchronisation. If every engineer turns to face the Product Owner or the manager when it’s their turn to speak, you don’t have a standup. You have a status update with better furniture.</p>
<p>The irony burns. We’ve taken a manifesto that explicitly values “responding to change over following a plan” and turned it into a rigid process where questioning the process is itself treated as a problem. That’s not Agile. That’s religion.</p>
<h2>“But, But… My Points!”</h2>
<p>I want to tell you about a conversation I had with one of my engineers. We were discussing a small change to how the team handled sprint carry-over. Nothing dramatic — just a suggestion that if a story isn’t 100% done, we shouldn’t count the percentage complete toward sprint completion. My reasoning was simple: optimise for <em>done</em>, not <em>half-done</em>. Finishing things matters more than starting them.</p>
<p>The engineer went quiet. You could see the calculation happening behind his eyes. Then, with genuine distress in his voice: “But, but… my points!”</p>
<p>He wasn’t being unreasonable. He was being perfectly rational within a broken system. We’d taught him — through months of velocity-driven conversations, through dashboards that tracked story points like a stock ticker, through sprint reviews where completion percentage was the headline number — that points were what mattered. Points were the scorecard. Points were his team’s identity.</p>
<p>The economist Charles Goodhart captured this decades before Scrum existed: “When a measure becomes a target, it ceases to be a good measure.” Story points were invented to help teams estimate complexity. The moment someone put them on a dashboard and started tracking velocity trends, they stopped being about estimation and started being about self-worth.</p>
<p>That engineer’s team had become a feature factory. They’d stopped caring about business outcomes entirely — the Product Owner measured success of ideas, and the team was there to do implementation. They’d lost ownership of results and replaced it with ownership of points. The output looked healthy. The outcome was invisible.</p>
<p>This is what ceremony without substance produces. Not failure — something worse. The <em>appearance</em> of success that prevents anyone from asking whether anything actually improved.</p>
<h2>Why the Cycle Repeats</h2>
<p>It’s easy to blame the certification industry. And to be fair, when a two-day Certified ScrumMaster course costs $1,500 per person and the Scrum Alliance has certified over a million professionals, there’s a self-perpetuating economy that has every incentive to defend the framework. Questioning Scrum threatens careers, certifications, and consultancies. The result is an immune system that attacks criticism.</p>
<p>But the real reasons go deeper.</p>
<p><strong>Management needs visibility, and process provides it.</strong> Tim Ottinger once described Scrum as “the box that XP comes in.” Allen Holub was more blunt in a conversation on Dave Farley’s Engineering Room: Scrum started out as a lightweight wrapper around Extreme Programming to make it palatable to management. That’s a perfect summary of what happened. Schwaber and Sutherland purposefully omitted XP’s technical practices — test-driven development, pair programming, continuous integration — to simplify organisational adoption. The engineering substance was stripped out so the process could be sold upward. And it worked — spectacularly well, from a sales perspective. The State of Agile survey found 66% of teams using Scrum but just 1% using XP. We kept the management wrapper and threw away the engineering practices it was supposed to protect.</p>
<p><strong>Organisations copy structure, not culture.</strong> When companies hear about Amazon’s two-pizza teams, they create small teams. When they hear about Spotify’s squads, they rename teams accordingly. But they don’t give those teams autonomy, because autonomy requires trust. They don’t invest in technical excellence, because that takes longer than renaming roles. They copy the visible structure and leave the invisible culture unchanged.</p>
<p>Bas Vodde put it brilliantly when discussing what happens when organisations try to scale Scrum: when you start to scale, some people have fear, and so they hire for positions like release train engineers to feel “safe.” That last word — safe — carries deliberate weight if you know what framework those roles come from.</p>
<p><strong>The practices that matter can’t be certified.</strong> Test-driven development, continuous integration, refactoring, pair programming, small batch delivery — these are the technical practices the manifesto authors identified as producing results. None of them require certification. None of them scale as consulting products. They require disciplined application over time. There’s no two-day course for TDD. Just years of practice.</p>
<p>This is why the cycle repeats: the things that work can’t be packaged. The things that can be packaged don’t work. The industry sells packages — and would you like a Jira subscription bundled with that?</p>
<h2>What the Successful Companies Actually Do</h2>
<p>Here’s the part that should frustrate you: the companies that ship at scale aren’t post-Agile. They’re pre-Agile. They’re doing what the manifesto described before anyone turned it into a product.</p>
<p>Amazon’s two-pizza teams own services end-to-end — from ideation to production operation. They don’t need another team’s permission to deploy. They don’t need a Scrum Master to facilitate communication. They choose their own practices. Notice what’s present: small teams, autonomy, ownership, accountability for outcomes. Notice what’s absent: mandated standups, story points, sprint ceremonies, certifications.</p>
<p>In 2012, Henrik Kniberg published a whitepaper describing how Spotify organised into squads, tribes, chapters, and guilds. The industry adopted it as a prescriptive framework. Companies worldwide renamed their teams “squads” and expected the results to follow. The irony is stunning: Spotify’s own Director of Engineering has said the model described in the whitepaper no longer reflects how Spotify works. They evolved past it. The whitepaper explicitly states that squads choose their own methodology. There is no mandated Scrum. The “Spotify model” is, at its core, an anti-model. The industry turned it into the very thing it was designed to prevent.</p>
<p>The common thread across these companies isn’t a framework. It’s a set of principles that were already in the Agile Manifesto: small autonomous cross-functional teams, technical excellence as non-negotiable, short feedback loops with real users, and ownership of outcomes rather than tickets. The manifesto authors would recognise it immediately. It’s what they were doing in 2001, before the certifications arrived.</p>
<h2>So Is Scrum the Problem?</h2>
<p>Let me be clear: Scrum isn’t fundamentally broken. I’ve seen it work. I’ve seen teams use it as a genuine tool for coordination, reflection, and incremental improvement. The framework itself — short iterations, inspect and adapt, cross-functional teams — embodies the manifesto’s spirit well enough.</p>
<p>The problem is how it’s sold. As a religion, not a tool. As a prescription, not a starting point. When your Scrum implementation has become a set of meetings with a two-week cadence where the standups don’t synchronise the team, the retros don’t produce change, and the velocity charts measure activity rather than achievement — you’re not doing Scrum badly. You’re doing Scrum <em>theatre</em>.</p>
<p>There’s a fascinating YouTube video from LEGO’s engineering team titled “Is SAFe Evil?” Their conclusion wasn’t that SAFe is inherently terrible — for them, it kind of worked. But it required careful adaptation, deep understanding of <em>why</em> the practices existed, and a willingness to throw out the parts that didn’t fit. In other words, they treated it as a framework to be adapted, not scripture to be followed. Which is, ironically, exactly what the Agile Manifesto asks you to do.</p>
<p>The critical distinction isn’t Scrum vs. Kanban vs. XP vs. nothing. It’s whether your process exists to serve the team or whether your team exists to serve the process. If you’ve ever heard someone justify a ceremony with “because we do Scrum,” you know which side of that line you’re standing on.</p>
<h2>The Blind Spot Nobody Talks About</h2>
<p>Here’s where Scrum genuinely fails, by design rather than implementation: it has no built-in mechanism for prioritising technical health.</p>
<p>When you put a Product Owner — typically a business person — in sole charge of the backlog, technical improvements compete directly with features. And they lose. Every time. Not because POs are malicious, but because the incentive structure makes it rational to defer technical work in favour of visible output. As I’ve written before, this is how you end up with codebases where deployment takes three days and nobody remembers why that one service needs to be restarted in a specific order on Tuesdays.</p>
<p>Here at Agoda, we’ve addressed this with guidance rather than mandate: 30% of engineering time should go toward technical improvements. It’s not a law of the land — it’s a health check. We measure it at the end of the quarter, and if it’s lower or higher we ask why. Sometimes there’s a good reason, and that’s fine. The point isn’t rigid enforcement. The point is that someone is asking the question. That alone changes behaviour more than any mandate could.</p>
<h2>Breaking the Cycle</h2>
<p>The answer isn’t another manifesto. It isn’t another framework. It isn’t renaming Agile to something that hasn’t been corrupted yet — Dave Thomas tried with “agility” and the industry will co-opt that too.</p>
<p>The answer is recognising that the split between “process” and “engineering” was always the problem. Agile tried to fix the process. It didn’t address the engineering. The result was a process movement that progressively excluded engineers — until, as Robert Martin observed, the programmers who started the movement no longer attended its conferences.</p>
<p>What comes next is what some of us have started calling product engineering. Not because the name matters — give it five years and someone will certify it — but because the concept matters. It treats engineering teams as product teams. Engineers don’t receive requirements — they discover problems. They don’t estimate stories — they measure outcomes. They don’t follow a process — they own a product.</p>
<p>There wasn’t a eureka moment for this. It was more of a slow recognition: this is what we do. It kind of works. It’s not perfect, but it works. And when we look at the companies that are genuinely successful, they’re doing similar things. The critical difference is that this isn’t something a process guide can give you. It’s a culture and mindset shift. It requires trust, autonomy, technical discipline, and a willingness to measure outcomes instead of velocity. You can’t buy that in a two-day course.</p>
<p>The Agile Manifesto was 68 words. It took thirty seconds to read. It took the industry twenty-five years and billions of dollars to misunderstand.</p>
<p>The practices that work haven’t changed: small teams, short feedback loops, technical excellence, customer focus, ownership. The manifesto got it right the first time. We just built a bamboo runway over the top of it and wondered why the planes never came.</p>
<p>Now, if you’ll excuse me, I need to go cancel a meeting that exists solely to discuss the outcomes of another meeting. But at least it’s in the Scrum Guide, so it must be working.</p>
]]></content:encoded>
  </item>
  <item>
    <title>The Lego Problem: Why Engineers Add Complexity Instead of Removing It</title>
    <link>https://blog.dicko.dev/posts/the-lego-problem-why-engineers-add-complexity-instead-of-removing-it/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/the-lego-problem-why-engineers-add-complexity-instead-of-removing-it/</guid>
    <pubDate>Mon, 02 Mar 2026 08:38:39 GMT</pubDate>
    <category>microservices</category>
    <category>software-architecture</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>programming</category>
    <description>Or: Why Your Architecture Reviews Have a Benefit Slide and No Cost Slide</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/the-lego-problem-why-engineers-add-complexity-instead-of-removing-it/cover.webp" alt=""></figure>
<p>It was 9:52 on a Wednesday morning and someone had drawn a new box on the whiteboard. The meeting room on level 7 still carried the stale warmth of the 10 AM standup that had run long, and the aircon hadn’t caught up yet. The box had a name — something with “orchestrator” in it — and three arrows pointing to existing services. The engineer presenting it was animated, enthusiastic, already two slides into the “how.” Nobody had asked “whether.” The product owner was nodding. The tech lead was sketching integration points on a napkin. And somewhere in the back of the room, a senior engineer was staring at the whiteboard with the quiet resignation of someone who knew they’d be on-call for this thing by March.</p>
<p>That box was never questioned. It was added, deployed, monitored, maintained, and is still running today. And the problem it was supposed to solve? It’s still there too — just with an extra component sitting on top of it.</p>
<p>This isn’t a post about microservices being bad. It’s not a post about simplicity as some kind of spiritual practice. It’s about a cognitive bias that’s been measured, replicated, and published in <em>Nature</em> — and that nobody in our industry talks about, despite it explaining half the architectural decisions that keep us up at night.</p>
<h2>The Experiment That Explains Your Architecture</h2>
<p>In 2021, researchers at the University of Virginia ran eight experiments with over 1,500 participants and found something that should be required reading for every engineering leader. When asked to improve something — anything — people systematically default to adding. Not because they considered subtraction and rejected it. They literally didn’t think of it.</p>
<p>The most striking experiment involved Lego. Participants were given a structure with a wobbly roof supported by a single off-centre pillar. The task: make it stable enough to hold a brick on top. Each added block cost 10 cents. Removing the pillar was free — the roof would rest flat on the base, perfectly stable, zero cost. But most people reached for more bricks.</p>
<p>As Antoine de Saint-Exupéry wrote in <em>Wind, Sand and Stars</em>, “Perfection is achieved, not when there is nothing more to add, but when there is nothing left to take away.” He was writing about aircraft design — another field where unnecessary complexity kills. And he published that in 1939, which means we’ve had nearly a century to absorb this lesson and we’re still adding pillars.</p>
<p>Now apply that to your last architecture review. Someone proposed a new service, a caching layer, a message queue. The discussion was about <em>how</em> to add it — API contracts, team ownership, deployment strategy. Subtraction never made it to the whiteboard. Not because anyone argued against it. Because nobody’s brain put it there.</p>
<h2>The Additive Default in Software</h2>
<p>You’ve sat in these meetings. We all have. The pattern is so consistent it’s almost formulaic.</p>
<p>“Performance is slow” → add a caching layer. “Services are coupled” → add a message queue. “Deployments are risky” → add a service mesh. “Auth is inconsistent” → add an API gateway.</p>
<p>Each addition sounds reasonable in isolation. Each one passes code review. Each one ships. And each one adds a component that can fail, an integration that needs monitoring, a system that needs maintaining — forever.</p>
<p>Here’s what gets discussed versus what doesn’t:</p>
<p>When you propose a new service, you talk about the API contract and team ownership. You don’t talk about the deployment pipeline, monitoring, on-call rotation, documentation, and inter-service testing that come with it. When you propose a caching layer, you talk about the performance improvement. You don’t talk about cache invalidation logic, consistency bugs, memory management, and cold-start problems. When you propose a message queue for decoupling, you don’t talk about ordering guarantees, dead letter handling, replay logic, and lag monitoring.</p>
<p>The pattern: every proposal has a benefit slide and no cost slide. Not because the proposer is lazy, but because the costs are invisible at proposal time and only become real at 3 AM six months later.</p>
<p>I wrote about the long-term decay this causes in <a href="https://medium.com/beer-and-servers-dont-mix/code-entropy-the-silent-killer-of-engineering-velocity">Code Entropy</a>. But entropy is the <em>effect</em>. Additive bias is the <em>cause</em>. Your brain is wired to make your codebase more complex, and it’s doing it without your permission.</p>
<h2>The Compound Cost Nobody Calculates</h2>
<p>This is where it gets worse than “more stuff = more problems.” Complexity compounds.</p>
<p>Two services with a message queue between them isn’t three components — it’s an exponential increase in failure modes. Service A can fail. Service B can fail. The queue can fail. A can fail while B is fine. B can fail while A is fine. The queue can lose messages. The queue can deliver duplicates. The queue can deliver out of order. Messages can poison the queue. A can produce faster than B consumes. And all of these interact with each other.</p>
<p>Architecture Components Integration Points Potential Failure Modes Monolith 1 0 Low 3 services 3 3+ Medium 10 services 10 20+ High 50 services 50 100+ Very High</p>
<p>Here’s the uncomfortable truth that most microservices advocates gloss over: <strong>microservices don’t reduce complexity. They redistribute it.</strong> From code complexity to operational complexity. From compile-time problems to runtime problems. From problems you find in development to problems you find in production. From problems your IDE catches to problems your <a href="https://blog.dicko.dev/posts/semantic-monitoring-the-question-youre-not-asking-about-your-production-systems/">monitoring dashboards</a> may or may not surface, depending on whether you’re measuring the right things.</p>
<p>The legendary basketball coach John Wooden once said, “Don’t mistake activity for achievement.” We’ve been remarkably active in our architecture reviews — adding services, adding layers, adding abstractions. The question is whether any of it achieved the simplicity we actually need to move fast and stay reliable.</p>
<h2>Kent Beck and the Forgotten Fourth Rule</h2>
<p>Most engineers know Kent Beck’s four rules of simple design. Or rather, they know three of them:</p>
<ul>
<li>Passes all tests- Reveals intention- No duplication- <strong>Fewest elements</strong>
That last rule — fewest elements — is the subtraction rule. It’s not “fewest <em>useful</em> elements” or “fewest elements <em>given our current architecture.</em>” It’s fewest. Period. If you can remove something without violating rules 1–3, you should.</li>
</ul>
<p>This connects to Dan McKinley’s influential essay <a href="https://boringtechnology.club/"><em>Choose Boring Technology</em></a> — the argument that every technology choice has an “innovation token” cost, and most organisations only have a few tokens to spend. Each addition to your stack spends a token whether you acknowledge it or not. That message queue you added? Token spent. That caching layer? Another token. That custom auth service? You’re out of tokens and you haven’t even started building the thing your users actually want.</p>
<h2>The Distributed Monolith: Adding Everything, Gaining Nothing</h2>
<p>The worst outcome of additive bias in microservices is something every architect has seen and nobody wants to admit they’ve built: the distributed monolith. You’ve added all the operational complexity of distributed systems while gaining none of the benefits. Services that must deploy together. Shared databases. Synchronous calls everywhere. No independent deployability.</p>
<p>You got the cost of microservices with the coupling of a monolith. This is what happens when the answer to every problem was “add a service” and the answer to no problem was “merge two services back together.”</p>
<p>Segment learned this the hard way. They grew to over 140 microservices, each with its own testing, deployment, and monitoring overhead. Three full-time engineers were spending most of their time just keeping the system alive instead of building features. They consolidated back into a monolith and saw immediate improvements in developer productivity — going from 32 improvements to shared libraries in a year to 46 improvements the year after consolidation. As their engineer Alexandra Noonan put it at QCon London, incorrectly implemented microservices can leave you unable to do product development because you’re drowning in the complexity.</p>
<p>This isn’t an anti-microservices argument. Microservices work if you do them right, and when you have a large enough scale of engineers, they’re genuinely needed. But the decision to add a service should require the same scrutiny as any other architectural decision — and that means asking the question we almost never ask.</p>
<h2>The Removal That Actually Helped</h2>
<p>Let me give you a real example. When we started building the new micro frontends for the YCS monolith split at Agoda, I made one decision early: no session state. All APIs would be stateless until someone could show me a concrete reason we needed it.</p>
<p>People were sceptical. Session state was just <em>how things were done</em>. The assumption was baked in so deeply that questioning it felt almost rude. But we held the line.</p>
<p>We spent two years migrating endpoints out. And nobody cared. It worked. We didn’t need state. We had a couple of cookies and that was it. Two years, hundreds of endpoints, and the subtraction of session state — something we’d been carrying around like a security blanket — caused zero problems while eliminating an entire category of complexity.</p>
<p>That’s what subtraction looks like in practice. Not dramatic. Not heroic. Just the quiet absence of a problem that never needed to exist.</p>
<p>And here’s the kicker: after two years, someone finally came to me with a legitimate use case for session state. We added it. And because we understood exactly why we were adding it and what problem it solved, we were able to drop our P99 server-side response time SLOs by whole seconds across multiple services. Whole seconds. Not milliseconds — seconds. That’s the difference between adding something because “we might need it” and adding something because you’ve proven you do. When you need it, you add it. Not before. YAGNI isn’t just a catchy acronym — it’s a strategy that pays compound interest when you finally do spend.</p>
<h2>Patch on Top of Patch</h2>
<p>Here’s the flip side — what happens when additive bias runs unchecked.</p>
<p>I was in a standup once where we were discussing a flow for onboarding new clients. There was a bug somewhere in the pipeline — about one in a hundred clients would end up with incorrect field data. The product owner asked if we could write a script to fix the bad data. Fair enough. We did — a simple SQL script, took no time at all.</p>
<p>Then someone said, “Now let’s run it on a schedule.”</p>
<p>I need you to sit with that for a moment. The proposed solution to a bug that corrupted one percent of client data was not to find and fix the bug. It was to run a script on a schedule that would periodically clean up after the bug. Patch on top of patch. A scheduled job to compensate for a defect that nobody wanted to dig into.</p>
<p>This is the stereotypical pattern I see everywhere. Rather than fixing root causes, we add another layer — a script, a service, a workaround, a retry mechanism, a compensating transaction. And then your system ends up as a tower of patches, each one depending on the others, none of them addressing the original problem. I’ve <a href="https://blog.dicko.dev/posts/the-prototype-that-became-production/">written about this pattern before</a> — how temporary things become permanent when nobody asks the next question.</p>
<p>E.F. Schumacher, the economist, wrote in <em>Small Is Beautiful</em>: “Any intelligent fool can make things bigger, more complex, and more violent. It takes a touch of genius — and a lot of courage — to move in the opposite direction.” That quote lands because it names the emotional truth: subtraction requires courage. Proposing to remove something means challenging a previous decision, possibly your own. It means saying “we got this wrong” or “this isn’t needed anymore.” That’s harder than proposing something new. And that’s exactly why it’s the more valuable skill.</p>
<h2>The Questions to Ask Before Adding Anything</h2>
<p>From handling large-scale failures over the years, I’ve learnt to love simplicity at scale. I don’t want smart things. I want dumb things that work reliably when everything else is on fire. The best design pattern at scale is simplicity. Not because it’s aesthetically pleasing — because it’s the only thing that survives contact with production at 3 AM.</p>
<p>I’ve been on the other side of this. Back in my previous life in event ticketing, when a major event went on sale at 9 AM, traffic would spike 100x in thirty seconds. Not gradually — thirty seconds. AWS autoscaling can’t respond that fast. I’ve watched it try. Your clever distributed caching strategy, your elegant circuit breakers, your beautifully abstracted service mesh — none of it matters when a hundred thousand people hit your system simultaneously and the infrastructure hasn’t caught up. The dumb things survived. The smart things fell over. As the military saying goes, no plan survives contact with the enemy — and neither does your architecture when the load arrives faster than your scaling policy.</p>
<p>Here are the questions that should be mandatory in every architecture review:</p>
<p><strong>“Can we solve this by removing something instead?”</strong> — Start here. Always. Your brain won’t suggest it naturally, so you need to force it onto the whiteboard.</p>
<p><strong>“What’s the simplest solution that works?”</strong> — Not the most elegant. Not the most scalable. Not the most “proper.” The question isn’t <a href="https://blog.dicko.dev/posts/the-question-that-changes-everything/">“how can we do this faster”</a> — it’s “how can we do this in a simpler way.”</p>
<p><strong>“What’s the cost of this addition over 5 years?”</strong> — Who maintains it? Who’s on-call? What happens when the person who built it leaves? Every component you add is a commitment your future team is making on your behalf.</p>
<p><strong>“What would happen if we did nothing?”</strong> — Sometimes the answer is “nothing bad.” That’s not laziness — that’s a signal.</p>
<p><strong>“Would I be comfortable explaining this in a post-mortem?”</strong> — If the answer is “we added it because it was interesting,” you have your answer.</p>
<p>The most valuable architectural skill isn’t knowing what to build — it’s knowing what to remove. And that skill is rare not because it’s difficult, but because it’s cognitively unnatural. You have to actively fight your own wiring to even consider it.</p>
<h2>The Bottom Line</h2>
<p>We quote “keep it simple” like a mantra, but we don’t live it. Sophistication in engineering culture means <em>more</em> patterns, <em>more</em> abstractions, <em>more</em> services. The engineer who proposes removing something is seen as less sophisticated than the one who proposes adding a new distributed transaction coordinator. We have it exactly backwards.</p>
<p>Your brain is working against you. The research is clear — you will default to adding, not removing, and you won’t even notice you’re doing it. The only defence is making subtraction a deliberate, first-class part of every architectural conversation. Ask the removal question first. Celebrate deleted code. Treat merging services back together as a sign of maturity, not defeat.</p>
<p>Microservices, caching layers, message queues, orchestrators — they’re all tools, and sometimes they’re the right tools. But every one of them should have to justify its existence against the alternative of doing less. And “doing less” should always be on the whiteboard.</p>
<p>Now, if you’ll excuse me, I need to go review an architecture proposal that adds three new services to solve a problem that might also be solved by deleting one. I already know which option requires more courage. I’m working on it.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Designing APIs That Outlive Your Current Sprint</title>
    <link>https://blog.dicko.dev/posts/designing-apis-that-outlive-your-current-sprint/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/designing-apis-that-outlive-your-current-sprint/</guid>
    <pubDate>Sun, 01 Mar 2026 06:46:50 GMT</pubDate>
    <category>api-design</category>
    <category>microservices</category>
    <category>software-engineering</category>
    <category>software-development</category>
    <category>backend-development</category>
    <description>Or: Why Your API Contract Is the Most Important Code You’ll Never Unit Test</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/designing-apis-that-outlive-your-current-sprint/cover.webp" alt=""></figure>
<p>It was 11:47 on a Tuesday morning when the war room on level 7 filled up. The Grafana dashboard on the NOC wall had gone from a lazy green river to an angry red waterfall — a massive spike in 4xx responses from the BFF, users hitting errors across multiple pages. Someone had already grabbed the Thai iced teas they’d never finish. The vinyl flooring squeaked under rolling chairs being pulled in too fast. Within ten minutes, we’d traced it to a BFF deployment: an intern had renamed a field from cust_id to customer_id. Reasonable change. Good naming, even. Their tests passed. Their code review passed. But every browser out there still running the previous version of the SPA — which was all of them — was parsing a field that no longer existed.</p>
<p>That’s not a story about an intern making a mistake. That’s a story about an organisation that treated an API like an implementation detail instead of what it actually is: a promise to every browser tab still open from this morning.</p>
<h2>The Promise Nobody Reads</h2>
<p>As the investor Warren Buffett once observed, <em>“It takes 20 years to build a reputation and five minutes to ruin it.”</em> He was talking about business, but he might as well have been describing your API contract. You spend months building trust with your consumers — they wire up their parsing, their error handling, their retry logic — and then one renamed field on a Tuesday morning means every user whose browser hasn’t fetched the latest JavaScript is now staring at a broken page.</p>
<p>Here’s the thing that took me years to internalise: your API is not your implementation. Your implementation is the messy, evolving, refactorable stuff behind the curtain. Your API is the curtain itself — the surface that every other team builds against, the contract that makes independent deployability possible. The moment you break that contract, you’re not deploying independently anymore. You’re coordinating. And coordination, as anyone who’s tried to get six teams to deploy in sequence on a Friday afternoon knows, is where velocity goes to die.</p>
<p>This post isn’t an API design textbook. It’s a set of practical checks — things you can run through in a code review in under ten minutes — that separate APIs that survive contact with the real world from APIs that generate a Slack message starting with “@channel urgent” every time someone deploys.</p>
<h2>The Intern, the BFF, and the SPA</h2>
<p>Let me give you the full picture of that war room incident, because the details matter more than the headline.</p>
<p>We had some interns working on one of our BFFs. The client code and the BFF lived in the same repo — a mono-repo setup that’s common at Agoda and plenty of other places. The intern looked at this and made a perfectly logical assumption: since the client and the server are in the same repo, they deploy together. So a breaking change to the API would be fine, right? The client code would just update at the same time.</p>
<p>They were young. It’s a reasonable mental model if you’ve never worked with SPAs before. A single-page application doesn’t redeploy when your server does. The JavaScript your users downloaded an hour ago — or yesterday, or last week if they haven’t refreshed — is still calling the old contract. Your server moved; your client didn’t. It’s the same repo, but it’s not the same deployment.</p>
<p>What makes this story worth telling isn’t the intern’s mistake — it’s that it was missed in review too. Experienced engineers looked at the diff, saw the improvement, and approved it. The breaking change was done the right way in every other sense: it was behind an experiment, which meant we didn’t need to roll back the deployment. We just killed the experiment. But the spike in 4xx responses was real, the war room was real, and the lesson was real.</p>
<p>The gap wasn’t knowledge. It was that we hadn’t built the instinct — the reflex — to ask: <em>“Can an existing client keep working without modifying their code?”</em> That single question, asked every time, is worth more than any API design document you’ll ever write.</p>
<h2>Your API Deserves the Same Thought as Your Architecture</h2>
<p>Here’s the uncomfortable truth about API design: you get one chance to get the contract right, and most teams spend more time debating variable names in a pull request than they do thinking about the shape of a response that three teams will depend on for the next three years.</p>
<p>We plan our database schemas. We whiteboard our system architecture. We argue about domain boundaries for weeks. But the API — the actual surface that consumers will build against, the thing that has to remain stable while everything behind it evolves — that gets designed in the time between standup and lunch by an engineer thinking about the feature they need to ship, not the contract they’re about to publish.</p>
<p>The physicist Richard Feynman once said, <em>“The first principle is that you must not fool yourself — and you are the easiest person to fool.”</em> We fool ourselves into thinking API design is just implementation that happens to be public. It isn’t. Implementation you can refactor on Tuesday. An API contract, once someone builds against it, is a conversation — and breaking changes are conversations you have at 2 AM on Slack with a very unhappy team.</p>
<p>Consistency is what makes an API learnable. When every endpoint follows the same conventions, engineers stop reading documentation and start predicting behaviour. They know where the resource will be, how errors will look, what happens when they send a bad request. That predictability isn’t a nice-to-have — it’s the difference between an API that people integrate with confidently and one that requires a Slack message to the owning team for every new endpoint. Getting this right upfront is dramatically cheaper than fixing it after consumers have already built against your inconsistencies.</p>
<p>So before we get into the specifics of what a good API contract looks like, let’s agree on one thing: this deserves deliberate thought. Not a design committee. Not a twelve-page RFC. Just the same care you’d give to any other architectural decision that will outlive your current sprint.</p>
<h2>Resource Naming: Let HTTP Do the Verbs</h2>
<p>The simplest API design rule has the biggest impact, and yet we keep getting it wrong. Let HTTP methods handle the verbs. Let your URLs describe the thing.</p>
<p>Don’t build GET /getBooking?id=123. Build GET /bookings/123. Don&#39;t build POST /createRatePlan. Build POST /rate-plans. The URL describes the resource; the method is the action. POST already means &quot;create.&quot; PUT already means &quot;update.&quot; You don&#39;t need to say it twice.</p>
<p>Use plural nouns for collections (/bookings, /rate-plans). Use hierarchy to show ownership (/properties/{id}/rate-plans). Pick a casing convention and stick with it — we&#39;re not going to relitigate camelCase versus kebab-case here, but inconsistency across endpoints is worse than either choice. Use query parameters for filtering and sorting, not for identity (/bookings?status=confirmed&amp;sort=checkIn).</p>
<p>If you want to see what “doing it right” looks like at scale, look at <a href="https://stripe.com/blog/api-versioning">Stripe’s API</a>. It’s almost universally cited as the gold standard in API design, and for good reason. Every resource follows the same conventions. Every endpoint behaves predictably. New engineers can guess the URL for an endpoint they’ve never used because the patterns are consistent. That didn’t happen by accident — it’s the result of a deliberate design review process that treats the API as a product, not a side effect of implementation. Brandur Leach wrote about their approach in detail back in 2017, and it’s still the benchmark a decade later.</p>
<p>One thing we’ve been historically good at here at Agoda — almost accidentally, if I’m being honest — is using business language in our APIs. If the team says “rate plan,” the URL says /rate-plans. If the domain calls it a &quot;booking,&quot; nobody&#39;s out there naming it /reservations because that&#39;s what they called it at their last company. This is the DDD concept of ubiquitous language applied to API design, and when it works, it eliminates an entire category of &quot;wait, what does this endpoint actually do?&quot; conversations.</p>
<h2>Versioning: You Probably Don’t Need It</h2>
<p>Most versioning debates I’ve sat through focus on the mechanism — URL path versus header — while ignoring the more important question: do you actually need a version bump at all?</p>
<p>Adding an optional request field with a sensible default? No version bump needed. Adding a new field to a response, assuming your consumers ignore unknown fields? No version bump. Adding an entirely new endpoint? Still no. But removing a field, renaming a field, changing a type, tightening validation that previously accepted something? Those are breaking changes. Those need a version or a deprecation plan or both.</p>
<p>The pattern is straightforward: additive changes are usually safe; subtractive or mutating changes are breaking. If you design for additive evolution from the start, you rarely need version bumps at all.</p>
<p>I always tell my engineers the same thing: <em>“Don’t do breaking changes.”</em> And they look at me like I’ve just asked them to write code without an IDE. So let me break this down. I contribute to <a href="https://nunit.org/">NUnit</a> when I’m not busy — which admittedly means it’s been a few years since my last PR — but NUnit went roughly nine years between major versions (version 3.0 shipped in 2015, version 4.0 in 2024). Now, NUnit is a library, not a BFF. But the principle is the same, just amplified: if an open-source project with thousands of consumers across the .NET ecosystem can go nine years without a breaking change, I’m fairly confident you can manage it with an internal API that has three consumers.</p>
<p>When you genuinely do need a version, start at /v1/. Evolve in place with additive changes. Reserve /v2/ for real, unavoidable breaking changes. Most services never need v2 if they were designed for evolution. Our APIs are about 90% internal — team-to-team — which gives us more flexibility than external-facing APIs. But &quot;internal&quot; doesn&#39;t mean &quot;casual.&quot; Internal consumers deserve the same stability as external ones. They just can&#39;t leave you a one-star review.</p>
<h2>The 200 OK That Lies to Your Face</h2>
<p>I need to talk about an anti-pattern that will resonate with every engineer who’s ever debugged a “successful” API call that actually failed.</p>
<p>By default, a Sangria GraphQL server will return HTTP 200 even when the query has execution errors, as long as the request itself was valid at the transport level. This follows the GraphQL spec convention, which treats GraphQL execution errors differently from transport errors. And on paper, that distinction makes sense. In practice, it means your monitoring thinks everything is fine, your load balancer thinks the service is healthy, your retry logic never triggers, and your engineers are left staring at a 200 OK response that contains, buried in the JSON, a field quietly announcing that everything went wrong.</p>
<p>The coach Vince Lombardi said, <em>“Practice does not make perfect. Only perfect practice makes perfect.”</em> We’ve been perfectly practising the wrong thing with our error contracts — returning success codes for failures so consistently that we’ve trained our entire monitoring infrastructure to miss real problems.</p>
<p>A good error response has three parts. The HTTP status code gives you the category — 4xx for client errors, 5xx for server errors. A stable, machine-readable error code gives your consumers something to program against — RATE_PLAN_NOT_AVAILABLE is infinitely more useful than parsing a free-form string that someone will reword next sprint and break your handling. And a human-readable message gives engineers and support something to read in logs.</p>
<pre><code>HTTP/1.1 422 Unprocessable Entity
</code></pre>
<pre><code>{
  &quot;error&quot;: {
    &quot;code&quot;: &quot;RATE_PLAN_NOT_AVAILABLE&quot;,
    &quot;message&quot;: &quot;The selected rate plan is no longer available for the requested dates.&quot;,
    &quot;details&quot;: {
      &quot;rate_plan_id&quot;: &quot;rp_2847&quot;,
      &quot;requested_check_in&quot;: &quot;2026-03-15&quot;
    }
  }
}
</code></pre>
<p>Compare that to what we’ve all seen in the wild:</p>
<pre><code>HTTP/1.1 200 OK
</code></pre>
<pre><code>{
  &quot;success&quot;: false,
  &quot;message&quot;: &quot;Something went wrong, please try again later.&quot;
}
</code></pre>
<p>The first one tells you what happened, gives your code something stable to switch on, and provides enough context to debug without digging through logs. The second one lies to your monitoring with a 200 status, gives your code nothing to work with, and tells the engineer exactly nothing about what went wrong. One error format per API. Documented. Treated as part of the contract. An undocumented error code is like an undocumented API endpoint — it exists, people depend on it, and it’ll break something when you change it.</p>
<h2>Backward Compatibility as a Discipline</h2>
<p>Everything above — naming, versioning, error contracts — serves one goal: your consumer can keep working without a code change when you deploy.</p>
<p>That’s it. That’s the test. Before every change, ask: <em>“Can an existing client keep working without modifying their code?”</em> If the answer is no, it’s a breaking change. Full stop. It needs a version bump, a deprecation plan, or a conversation.</p>
<p>The deprecation workflow we follow isn’t glamorous, but it works. Announce the deprecation with a timeline — “v1 sunset in 90 days.” Add a Deprecation header so consumers can detect it programmatically. Provide a migration guide that&#39;s actually useful, not a changelog formatted as a guide. Monitor usage — reach out to teams still on the old version. Then, and only then, sunset it.</p>
<p>The architect Christopher Alexander wrote, <em>“When you build a thing you cannot merely build that thing in isolation, but must repair the world around it, and within it, so that the larger world at that one place becomes more coherent.”</em> We’ve been doing the opposite with our APIs — building things that make the world around them less coherent, one breaking change at a time. Every field rename, every tightened validation, every removed endpoint is a small fracture in the trust that lets teams deploy independently.</p>
<p>If you’ve read my post on <a href="https://blog.dicko.dev/posts/the-prototype-that-became-production/">backwards compatibility testing patterns</a>, you’ll know we’ve built tooling at Agoda to catch these issues before they escape — compiling against both the new interfaces and the last published version to make the build system enforce what code review sometimes misses. That’s the safety net. But the real discipline happens earlier, in design, when you decide to add an optional field with a default instead of renaming the existing one.</p>
<h2>The Five-Minute Code Review Checklist</h2>
<p>If I’ve done my job with this post, you should be able to run through these checks on any API change in under ten minutes:</p>
<p>Does the URL describe a resource with a plural noun, or is it a verb disguised as an endpoint? Is this an additive change, or does it remove, rename, or mutate something existing? If it’s breaking, is there a versioning and deprecation plan? Does the error response use proper HTTP status codes, or are we hiding failures behind 200 OK? Can an existing client — including one running JavaScript downloaded three days ago — keep working without a code change?</p>
<p>That last question is the only one that really matters. The rest are just ways of getting there.</p>
<h2>The Bottom Line</h2>
<p>The teams that build APIs worth depending on aren’t doing anything magical. They’re asking one question, consistently, at every boundary: <em>“Can my consumers keep working when I deploy?”</em> They’re treating their API as a product with its own design standards, not as a side effect of their implementation. They’re designing for additive evolution so they rarely need breaking changes. And when they do need to break something, they’re giving consumers a migration path that respects their time.</p>
<p>As I’ve <a href="https://blog.dicko.dev/posts/6-golden-rules-for-library-development-ensuring-stability-and-reliability/">written before</a>, the discipline of library APIs and service APIs is the same discipline — keep your promises, evolve without breaking, and treat every public surface as a contract that someone’s production system depends on. Your <a href="https://blog.dicko.dev/posts/technical-debt-the-elephant-in-the-scrum-room/">code entropy</a> grows fastest at the boundaries between systems, and APIs are the biggest boundaries you have.</p>
<p>Now, if you’ll excuse me, I need to go review a PR where someone’s proposing to “just rename” a response field on one of our BFFs. The field name is genuinely terrible — I think it was named during a Friday afternoon session that may or may not have involved beer — but there are browser tabs out there right now parsing that field, and every one of them is a promise I intend to keep.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Show Me the Architecture and I’ll Show You the Silos</title>
    <link>https://blog.dicko.dev/posts/show-me-the-architecture-and-ill-show-you-the-silos/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/show-me-the-architecture-and-ill-show-you-the-silos/</guid>
    <pubDate>Sun, 01 Mar 2026 06:43:10 GMT</pubDate>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>microservices</category>
    <category>engineering-leadership</category>
    <category>software-architecture</category>
    <description>Or: How We Broke Down Our Monolith and Accidentally Built Walls Between Our Teams</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/show-me-the-architecture-and-ill-show-you-the-silos/cover.webp" alt=""></figure>
<p>Somchai’s cursor hovered over the merge button for three seconds too long. It was 4:20 on a Wednesday, the afternoon light cutting low through the floor-to-ceiling windows on level 27, turning everyone’s monitors into mirrors. He’d just spent the better part of a week implementing client-side round-robin load balancing for his team’s micro frontend — a clean solution, well-tested, exactly the kind of thing you’d want a senior engineer to ship. The problem was that two floors down, another team had shipped an almost identical implementation six weeks earlier. And on the floor above, a third team had built their own version in Scala instead of C#. Nobody had done anything wrong. That was the part that sat in my stomach like cold rice. Three teams, three implementations, zero communication failures — because there was nothing in the architecture that made communication necessary anymore.</p>
<p>The monolith was dead, and we’d thrown a party. We just hadn’t noticed what we’d buried with it.</p>
<p>Charlie Munger, the late investor and Berkshire Hathaway vice chairman, had a line that should be written in permanent marker on every engineering leader’s whiteboard: “Show me the incentives, and I’ll show you the outcome.” He was talking about financial markets, but he might as well have been describing what happens when you decompose a monolithic codebase into microservices and micro frontends. You change the architecture, you change the incentive structure, and the outcomes follow with the reliability of gravity. We just don’t always like what those outcomes turn out to be.</p>
<h2>The Monolith Was Your Worst Collaboration Tool (and Your Best)</h2>
<p>Here’s the uncomfortable truth that nobody mentions in the “monolith to microservices” migration talks at conferences: the monolith was forcing your teams to collaborate, and it was doing it for free.</p>
<p>Not because anyone designed it that way. Not because there was a collaboration strategy or a cross-team engagement framework or whatever your favourite consultancy is selling this quarter. The monolith forced collaboration because the code lived in the same place. Your changes touched other people’s code. Their deployments included your features. The shared pain of a broken build at 5 PM on a Thursday was everyone’s pain, felt simultaneously, impossible to ignore.</p>
<p>You didn’t choose to learn what the team down the hall was building. You couldn’t avoid it. Their new caching layer showed up in your pull request diff. Their database migration broke your integration tests. The coupling was a constraint, yes, and it slowed you down, absolutely — but it also meant that when someone built a round-robin implementation, everyone saw it. When someone introduced a new pattern, it was visible in the same repository you opened every morning.</p>
<p>Then you did the right thing. You decomposed the monolith. You gave teams autonomy, independent deployments, their own repositories, their own CI pipelines. And every single one of those improvements — each of which was genuinely, defensibly correct — removed a forcing function for collaboration that you didn’t know you were depending on.</p>
<p>The urban planner Jane Jacobs spent her career arguing that the physical design of cities shapes how people interact. Short blocks and mixed-use streets create what she called “the ballet of the sidewalk” — those constant, casual, seemingly random encounters between strangers that form the foundation of a neighbourhood’s trust and social fabric. Long blocks, highways, and separated zones kill that ballet. Not through malice. Through design.</p>
<p>Your monolith was a short city block. Everyone bumped into each other. Your microservices architecture is a suburban highway system — efficient for getting from A to B, but don’t expect your neighbours to know your name.</p>
<h2>The Redux Moment</h2>
<p>The moment I knew we had a problem wasn’t the round-robin duplication. It was earlier, and subtler.</p>
<p>One team removed Redux from their micro frontend and introduced a different state management approach. Taken in isolation, this was a perfectly good decision. Even Dan Abramov, the creator of Redux, has publicly stated he wouldn’t recommend it for new projects. The team evaluated their options, picked something that better fit their needs, and shipped it. Textbook engineering autonomy.</p>
<p>But they did it for themselves. Just their micro frontend. And suddenly you’ve got a divergence event. One team is on a fundamentally different state management paradigm than the other nine. An engineer who’s been contributing across multiple micro frontends now needs to context-switch between two mental models. Cross-contribution just got harder. Consistency just took a hit.</p>
<p>A friend of mine went to work at Facebook, where the backend runs on Hack — their custom fork of PHP. I asked him the obvious question: “So, is the code shit?” He didn’t even pause. “Yes, it’s shit. But it’s consistently shit, and that makes it okay.” He wasn’t joking. Consistency at scale is worth more than local excellence, because consistency means any engineer can move between systems without relearning the fundamentals. The moment you lose that, every team boundary becomes a context-switch tax that compounds with every rotation, every cross-team contribution, every production incident that requires someone unfamiliar to jump in.</p>
<p>And here’s the thing that makes this genuinely difficult: nobody did anything wrong. The team exercised the autonomy that the architecture gave them. The decision was technically sound. The problem isn’t that they made a bad choice — it’s that the architecture made the choice invisible to everyone else until it was already shipped.</p>
<p>In a monolith, someone would have seen the pull request. Someone would have said, “Interesting — should we do this everywhere, or are we deliberately keeping both?” The shared codebase would have made the divergence a conversation before it became a fact.</p>
<p>Some divergence is innovation. One team finds a better approach and the rest of the organisation benefits. Five teams independently building the same utility is waste. The leader’s job isn’t to prevent divergence — it’s to maintain enough connective tissue that teams can see each other’s work and consciously choose to diverge rather than accidentally duplicate.</p>
<h2>Conway Had a Law. We Proved It.</h2>
<p>There’s a reason Melvin Conway’s fifty-year-old observation keeps showing up in architecture talks: because organisations keep rediscovering it the hard way. Your system’s structure will mirror your organisation’s communication structure. Or, to flip it around: show me the architecture, and I’ll show you where the silos will form.</p>
<p>When we had a monolith, everyone communicated because the code demanded it. When we moved to microservices with team-aligned boundaries, the teams optimised locally — exactly as Conway’s Law predicts — because that’s what the architecture rewarded. You can’t blame a team for focusing on their own service when the architecture, the CI pipeline, the deployment model, and the code review process all reinforce that boundary.</p>
<p>This is the key insight that separates the conversation from every other “break down silos” article you’ve read: the decomposition that made our architecture better made our collaboration worse, and both of those things were entirely predictable. We traded one kind of coupling — code coupling, which forced uncomfortable but valuable interactions — for another kind of problem: divergence, duplication, and invisible decision-making. And unlike the monolith, microservices don’t have a built-in forcing function for alignment.</p>
<p>So the post isn’t “silos are bad, go fix your culture.” Culture doesn’t exist in a vacuum. Culture is the emergent behaviour of the systems you’ve built. The real message is this: when you remove the architectural forcing function, you have to build an intentional one, or entropy wins.</p>
<h2>Building the Intentional Forcing Functions</h2>
<p>If the architecture won’t force collaboration for you anymore, you need to create the mechanisms deliberately. Here’s what we’ve tried, what worked, and what fell flat on its face.</p>
<h2>Platform Teams as Enabling Teams, Not Gatekeepers</h2>
<p>The reflexive response to divergence is centralisation — create a platform team, make everyone go through them, problem solved. Except it’s not, because you’ve just created a bottleneck that slows everyone down and teaches them to work around the system instead of with it.</p>
<p>We had this exact problem with an experimentation library owned by a platform team. Every team needed to run A/B tests, and the library was the official way to do it. But when teams needed changes or extensions, the queue was long and the platform team had their own priorities. So teams didn’t wait. They built on top. Wrappers, helper functions, custom abstractions layered over the library — reasonable decisions made independently by reasonable people. And then the wrappers got copied and pasted between repositories, each copy diverging slightly as different teams modified them for their own needs. Within a year, you had a dozen variations of experimentation logic scattered across the codebase, none of them quite the same, all of them one library update away from breaking. You wanted consistency; you got duplication plus fragility.</p>
<p>The Team Topologies model gets this right: a platform team should be an enabling team, not a gatekeeper. The distinction is critical. A platform team that requires pull requests creates a bottleneck. A platform team that provides golden paths — self-service tooling, shared libraries, well-documented patterns — creates alignment without coupling. The round-robin problem doesn’t get solved by mandating “everyone must use the approved implementation.” It gets solved by making the approved implementation so easy to adopt that building your own would be harder.</p>
<h2>Mob Sessions: Force It, Then Let It Breathe</h2>
<p>We introduced cross-team mob programming sessions — each participating team sends a representative every two weeks for a half-day session to work together on a common problem. And I’ll be honest with you: we had to mandate participation at first. Not everyone wanted to be there. The early sessions felt forced because they were forced.</p>
<p>That’s okay. As a manager, you can tell your people what to do. You shouldn’t make too much of a habit of it — over-directing is ownership-corrosive — but sometimes the initial push is necessary to start a habit that eventually sustains itself.</p>
<p>The key lessons: attend the first sessions yourself to set expectations for how mob programming works, who drives, how you rotate, what outcomes you’re aiming for. Get regular feedback. Adjust the format ruthlessly based on what you hear. The original mandatory format eventually tanked — people hated being forced into sessions on topics irrelevant to their work. When we evolved to opt-in, interest-based labs where engineers self-selected into problems they cared about, the dynamic completely changed. The forcing function got the ball rolling; the format evolution kept it alive.</p>
<h2>Combined Sprint Demos: Make Divergence Visible</h2>
<p>This is the one I’d recommend to anyone running multiple teams, and it costs almost nothing. Once a month, all ten teams present their sprint demos to each other in a single session.</p>
<p>The Redux moment I described earlier? In a combined demo, it becomes visible immediately. One team says “we replaced state management” and nine other teams hear it. Half the room is already nodding because they’ve been quietly hating Redux for years. Someone asks “should we all do this?” The conversation happens before the divergence becomes entrenched.</p>
<p>The demo doesn’t prevent divergence. It makes divergence visible and deliberate. That’s the difference between innovation and accidental duplication. Innovation is when you diverge consciously, knowing what everyone else is doing, and choosing a different path for good reasons. Duplication is when you diverge because you had no idea anyone else had already solved the problem.</p>
<h2>Cross-Team Technical Work as an Expectation</h2>
<p>Your staff engineers and tech leads should be spending meaningful time outside their own team’s sprint. I’ve <a href="https://claude.ai/chat/LINK_TO_IMPACT_POST">written about this at length before</a> — your impact needs to extend beyond your team’s backlog. If an IC4+ in a medium-to-large engineering organisation is only working within their team’s sprint boundaries, two things are probably true — they’re not having the wider impact their level demands, and they’re going to get bored and leave.</p>
<p>We back this up with weak code ownership. Following Martin Fowler’s model: every module has an owner, but anyone can contribute via pull request. The owner reviews and maintains quality, but the doors are open. If your team needs something changed in another team’s service, the expectation is that you open the PR yourself rather than filing a request and waiting. This has a compounding effect — it builds cross-team knowledge, it removes bottlenecks, and it moves you closer to genuine feature teams where ownership follows the customer problem rather than the technical component. Tools like Sourcegraph have been a godsend here — its batch changes feature lets you open pull requests across dozens of repositories simultaneously, which means a single engineer can drive a library upgrade or a pattern migration across the entire microservices estate in an afternoon rather than filing tickets with fifteen teams and waiting three sprints.</p>
<h2>Give People Time and Make It Visible</h2>
<p>None of this works if engineers don’t have protected time for cross-team work. Atlassian gives individuals 20% time. Here at Agoda, we use <a href="https://blog.dicko.dev/posts/technical-debt-the-elephant-in-the-scrum-room/">30% of engineering time as a guideline</a> for technical improvements — it’s not mandatory, it’s not documented in a policy, it’s guidance that teams balance against their own priorities. But the expectation that a meaningful chunk of engineering time lives outside feature delivery is explicit, and that time includes cross-team initiatives.</p>
<p>But time alone isn’t enough. You need visibility from the top down. Get leadership and business support for the idea that collaboration across team boundaries isn’t a distraction from “real work” — it is the work. Facilitate the infrastructure: regular tech talks, training sessions, brown bag lunches. Have budget, have rooms, have people who help engineers organise these things. Don’t create administrative barriers to collaboration, and if barriers exist, tear them down.</p>
<h2>Shared Goals: The Foundation Everything Else Sits On</h2>
<p>If I had to pick one forcing function above all others, it’s this: shared goals at the department level. As a tech department, what do we collectively want to achieve this quarter or this year? What do all of our teams value? What do we agree is the most important thing for moving us forward?</p>
<p>Once you’ve agreed on the what, set the measurement, identify the teams that will contribute, and create the structures for those teams to work together toward a shared long-term outcome. When ten teams are all measured against the same result, collaboration stops being a nice-to-have and starts being the rational behaviour. You’ve aligned the incentive structure with the collaboration you actually want.</p>
<p>Munger was right. Show me the incentives, and I’ll show you the outcome. Show me shared goals backed by shared measurement, and I’ll show you teams that collaborate without being told to.</p>
<h2>Before You Say “Modular Monolith”</h2>
<p>I can already hear it. Someone is reading this and thinking: “This is exactly why modular monoliths are the answer.” It’s the hot take of the moment, and I understand the appeal — you get the deployment simplicity of a monolith with the logical separation of microservices. Best of both worlds.</p>
<p>Except module boundaries are only a single step away from separate git repositories. The moment teams start owning modules, the same silo dynamics appear inside the monolith that we saw outside it. Teams stop looking at each other’s modules. Knowledge clusters around ownership boundaries. Someone builds a utility in their module that three other modules could use, but nobody knows it exists because nobody’s browsing code outside their own area anymore.</p>
<p>We’ve seen it. The modular monolith doesn’t solve the collaboration problem — it just changes the radius. If you don’t build intentional forcing functions, the architecture will teach your teams to optimise locally regardless of whether “locally” means a microservice, a module, or a folder in a monorepo. The shape of the silo changes. The silo doesn’t.</p>
<h2>The Architecture Is Not Morally Superior</h2>
<p>You’re not going back to the monolith. Nobody wants to. At scale, and done the right way, the microservices architecture is better in a hundred ways that matter: independent deployments, team autonomy, technology flexibility, fault isolation. These are real, significant advantages, and I wouldn’t trade them. At a smaller scale, the monolith is almost certainly the better choice — but that’s a different post.</p>
<p>But you have to be honest about what you lost when you left the monolith behind. The monolith forced collaboration as a side effect. Microservices force isolation as a side effect. Neither architecture is morally superior — they just have different default behaviours, and those defaults shape your teams in ways that are entirely predictable if you’re paying attention.</p>
<p>The question isn’t which architecture to choose. You’ve already chosen. The question is whether you’re willing to build the intentional systems that compensate for whichever architecture you chose.</p>
<p>Show me the architecture, and I’ll show you where the silos will form. Show me the forcing functions, and I’ll show you whether the leader knew they were coming.</p>
<p>Now, if you’ll excuse me, I need to go check how many implementations of client-side retry logic we have across our micro frontends. I’m afraid to count, but I’m more afraid not to.</p>
]]></content:encoded>
  </item>
  <item>
    <title>Resilience Isn’t a Feature — It’s the Absence of Naivety</title>
    <link>https://blog.dicko.dev/posts/resilience-isnt-a-feature-its-the-absence-of-naivety/</link>
    <guid isPermaLink="true">https://blog.dicko.dev/posts/resilience-isnt-a-feature-its-the-absence-of-naivety/</guid>
    <pubDate>Sat, 28 Feb 2026 05:50:47 GMT</pubDate>
    <category>microservices</category>
    <category>software-development</category>
    <category>software-engineering</category>
    <category>resilience</category>
    <category>distributed-systems</category>
    <description>Or: Why Your System Is One Slow Query Away from a Very Bad Night</description>
    <content:encoded><![CDATA[<figure><img src="https://blog.dicko.dev/posts/resilience-isnt-a-feature-its-the-absence-of-naivety/cover.webp" alt=""></figure>
<p>The Grafana dashboard had been green all day. At 2:14 AM on a Thursday, the on-call engineer’s phone lit up the bedside table — not one alert, but eleven, arriving in a cascade that sounded like a slot machine paying out in the worst way possible. He sat up, squinted at the screen, and watched three services go red in the time it took to unlock his laptop. The vinyl floor was cold under his feet as he walked to the kitchen for water he wouldn’t drink. By the time he’d opened the VPN, the fourth service was down.</p>
<p>That’s not a hypothetical. That’s what happens when distributed systems discover, at the worst possible moment, that nobody taught them how to fail.</p>
<p>Every microservices tutorial on the internet will teach you how to make services talk to each other. Almost none of them cover what happens when they stop. This is like teaching someone to drive by explaining the accelerator and the steering wheel but skipping the brakes, the seatbelt, and what to do when the engine catches fire on the expressway.</p>
<p>As the management theorist Peter Drucker put it, <em>“The greatest danger in times of turbulence is not the turbulence — it is to act with yesterday’s logic.”</em> In distributed systems, yesterday’s logic is the assumption that your dependencies will be there when you need them. Today’s reality is that at any given moment in a system with fifty services, something is slow, something is timing out, something is returning garbage data, and something is completely down. The question isn’t whether your system handles failure. It’s whether it handles failure gracefully — or whether it handles it by propagating the problem to everything else until the whole thing collapses like dominoes on a wobbly table.</p>
<h2>The Anatomy of a Cascade</h2>
<p>Let’s walk through what actually happens, because understanding the mechanics is the first step toward building the instinct to prevent it.</p>
<p>Service A calls Service B. Service B calls Service C. Service C calls a database. The database is slow — maybe someone ran an analytics query, maybe a table lock went sideways, maybe it’s Tuesday and Tuesdays are just like that sometimes. The query that normally returns in 40 milliseconds is now taking 90 seconds.</p>
<p>Service C’s threads are all waiting. Its thread pool exhausts. Service B’s threads are all waiting on Service C. Its pool exhausts. Service A’s threads are all waiting on Service B. Its pool exhausts. The load balancer starts returning 503s. Users see error pages. Your monitoring lights up like a Christmas tree. And somewhere in the middle of it all, a single slow database query is sitting there, blissfully unaware that it just brought down three services and ruined several thousand people’s evening.</p>
<blockquote>
<p>“For want of a nail the shoe was lost. For want of a shoe the horse was lost. For want of a horse the rider was lost.”* — Proverb, dating to the 13th century*</p>
</blockquote>
<p>That isn’t a metaphor. That is literally how cascading failures work. A missing nail — a single slow dependency — collapses the entire chain because nothing in it knows how to fail gracefully.</p>
<h2>The Three Things Every External Call Needs</h2>
<p>We’re not talking about advanced patterns here. We’re talking about the basics — the seatbelts and brakes that should be on every service before it goes anywhere near production.</p>
<p><strong>Timeouts: Stop Waiting for Godot</strong></p>
<p>Without a timeout, a slow dependency turns your fast service into a slow service. Your threads block. Your pool exhausts. Your callers block. The cascade begins.</p>
<p>Every external call needs a timeout. HTTP calls. Database queries. gRPC calls. Cache lookups. Message queue operations. No exceptions.</p>
<p>Here’s the thing that still surprises me after years of doing this: the default timeouts on most HTTP and SQL clients are insane. We’re talking minutes. <em>Minutes.</em> In what universe does waiting two minutes for a response represent a functioning system? At Agoda, the first thing we change out of the box is setting timeouts to milliseconds. If you haven’t responded in less than a second, something is wrong, and I’d rather fail fast, forget about it, and retry the next node than sit there holding the line while the whole system backs up behind me.</p>
<p>A timeout is actually an optimistic pattern, if you think about it. It says “I believe this call will succeed — but if it doesn’t respond quickly, something is wrong and I’d rather know now.” Without a timeout, you’re saying “I’ll wait as long as it takes,” which in a distributed system can mean forever. And forever is a long time to block a thread.</p>
<p><strong>Retries: The Second Chance (Done Right)</strong></p>
<p>The naive approach — retry immediately on failure — is the pattern that turns a struggling service into a dead service. If a service is slow because it’s overloaded, hammering it with retries is like seeing someone drowning and throwing them more water.</p>
<p>We learned this the hard way at Agoda. We had an incident where a service five levels deep in the call stack started timing out. Every layer above it had retries configured. The web browser had retries. The BFF had retries. The mid-tier services had retries. When that bottom service got slow, it didn’t get fewer requests to help it recover — it got exponentially <em>more</em> requests as every layer above it dutifully retried. The system that was under pressure got crushed under the weight of everyone trying to help.</p>
<p>After that incident, we started adding and propagating headers with retry counts so that downstream systems could see how many times a request had already been retried and react accordingly. If you’re the fifth retry of a request that’s already been retried by three layers above you, maybe the kindest thing you can do is say no.</p>
<p>The evolution of retry strategies tells you everything about how we’ve learned to respect distributed systems:</p>
<p>No retry means transient errors cause unnecessary failures. Immediate retry means you hammer the dying service. Exponential backoff gives the service time to recover. Backoff with jitter prevents the thundering herd — that terrifying phenomenon where a thousand clients all retry at exactly the same backoff interval, creating a synchronised retry storm that’s worse than the original load.</p>
<p><strong>Circuit Breakers: Knowing When to Stop Trying</strong></p>
<p>Your house has a circuit breaker. When the electrical current exceeds safe levels, the breaker trips and cuts the circuit. It doesn’t keep pushing current through and hope the wiring doesn’t catch fire. It stops and waits.</p>
<p>Software circuit breakers work the same way. When a dependency’s failure rate exceeds a threshold, the circuit opens and requests fail immediately — you don’t even try. After a timeout period, you let a few requests through to see if the service has recovered. If it has, the circuit closes and normal operation resumes. If it hasn’t, the circuit stays open.</p>
<p>Here’s the emotional insight that most engineers miss: failing fast is an act of kindness to the rest of the system. By <em>not</em> calling a broken service, you’re protecting your own resources, protecting the broken service from additional load it can’t handle, and giving your users a fast error instead of a slow timeout. A quick “no” is almost always better than a long silence.</p>
<h2>Bulkheads: Containing the Blast Radius</h2>
<p>The name comes from the watertight compartments in ships. The Titanic sank because water flowed between compartments. In software, without bulkheads, one failing dependency drowns your entire thread pool.</p>
<p>Without isolation, all your threads live in one pool. A slow call to Service B consumes threads. Perfectly healthy calls to Service A can’t get threads because they’re all stuck waiting on Service B. With bulkheads, each dependency gets its own pool. Service B’s pool exhausts, but Service A’s pool is untouched and keeps working.</p>
<p>Netflix built their entire Hystrix library because of exactly this problem. When one of their hundreds of services slowed down, it would consume the thread pools of its callers, which consumed the thread pools of <em>their</em> callers, until the entire system collapsed. Hystrix introduced bulkheads, circuit breakers, and fallbacks as first-class architectural concerns — not afterthoughts bolted on after the third outage.</p>
<p>At Agoda, we took a different path to the same destination. Over a decade ago, we had a massive outage caused by a load balancer failure. That experience burned us deeply enough that we removed traditional load balancers entirely and moved to client-side round robin with service discovery. Our teams wrote custom HTTP and SQL clients that handle resilience at the client level — service discovery, round robin, retries, timeouts, all baked in. Not only did this eliminate that single point of failure, it saved us significant infrastructure cost. These libraries are community-maintained across the engineering organisation, and we’ve even open-sourced some of them.</p>
<p>Recently, we’ve been moving toward Envoy for some of this, but the principle remains: resilience lives in the client, not in a shared piece of infrastructure that becomes everyone’s problem when it goes down.</p>
<h2>Fallbacks: The Business Decision Nobody Made</h2>
<p>This is where the conversation shifts from engineering patterns to leadership responsibility, and where I need every engineering manager and product owner reading this to pay attention.</p>
<p>When a service is down, you have options. You can return cached data — show yesterday’s prices. You can return a default — “Price unavailable, call for a quote.” You can degrade gracefully — hide the reviews section entirely. Or you can fail fast and show an error page.</p>
<p>Here’s the problem: the choice between these options is a business decision, not an engineering decision. Should you show stale prices or no prices? If you’re running a hotel booking site, showing yesterday’s price might lead to a booking at the wrong rate — that’s a financial cost. But showing no price means the user leaves — that’s a revenue cost. Which is worse?</p>
<p>Your engineer should not be making this call alone at 3 AM with one eye open and cold coffee going stale on the desk. Your product owner should have made it at 3 PM last Wednesday, and it should be documented and implemented as a fallback strategy <em>before</em> the failure happens.</p>
<blockquote>
<p>“By failing to prepare, you are preparing to fail.”* — Benjamin Franklin*</p>
</blockquote>
<p>Franklin wasn’t thinking about microservices, but he nailed it anyway. Resilience isn’t something you bolt on during the incident. It’s a conversation you have before the incident, between product and engineering, about what “good enough” looks like when “perfect” isn’t available.</p>
<p>Something we do at Agoda that I think is genuinely unique is feature shedding. When traffic spikes beyond what the system can comfortably handle, we start deliberately dropping compute-expensive features. Not randomly — strategically. We identify which features are expensive to render and which are critical to the core user journey, and when the system is under pressure, we shed the expensive non-critical ones. The user still gets a functioning booking experience. They just might not get personalised recommendations or that fancy interactive map while the system is under stress. It’s the difference between dimming the lights to keep the power on and letting the whole grid go dark.</p>
<h2>The Real Cost of Not Being Resilient</h2>
<p>Let’s do the maths, because this is where abstract patterns become concrete urgency.</p>
<p>Service A calls Service B calls Service C. Without timeouts, if C takes 60 seconds, B waits 60 seconds, A waits 60 seconds plus B’s processing time. If A has 100 threads and C goes slow, all 100 threads block within minutes. Service A is now effectively down — not because anything is wrong with A’s code, but because it trusted C unconditionally.</p>
<p>Amazon’s 2017 S3 outage is the canonical example. A simple operational error in S3 caused failures across dozens of AWS services because so many services depended on S3 without adequate resilience patterns. The blast radius extended far beyond what anyone anticipated because the dependency was so deeply embedded in everything. One service. Hundreds of dependents. Hours of cascading failure.</p>
<p>Michael Nygard’s book <em>Release It!</em> frames this perfectly: stability patterns exist because instability patterns are the default. In a distributed system, the absence of resilience patterns isn’t a neutral state. It’s an active risk. You’re not being conservative by skipping the circuit breaker. You’re being reckless.</p>
<h2>Be Your Own Chaos</h2>
<p>You might be wondering whether we do chaos engineering at Agoda — intentionally breaking things to test resilience. The honest answer is: not formally. We are our own chaos.</p>
<p>With a culture of ruthless experimentation and thousands of A/B tests running simultaneously, we create enough organic turbulence to keep our systems honest. We do scheduled traffic shifts to test the capacity of our fallback paths, and the feature shedding I mentioned earlier gets exercised more often than you’d think. But we haven’t needed to inject artificial chaos because the combination of scale, experimentation, and continuous deployment provides a steady stream of real-world failure scenarios.</p>
<p>That said, this isn’t a recommendation to skip chaos engineering. It’s an observation that if you’re running at sufficient scale with sufficient experimentation velocity, your system is already being stress-tested constantly. The question is whether you’re paying attention to what those tests are telling you.</p>
<h2>Resilience as Culture, Not Configuration</h2>
<p>Here’s where we land, and it’s the point I want to leave you with: resilience isn’t a sprint feature. You don’t ship it in a two-week cycle and move on. It’s a design posture — a default way of thinking about every external call, every dependency, every integration point.</p>
<p>The engineers who build resilient systems aren’t the ones who memorised the patterns. They’re the ones who <em>assume</em> every external call will fail and design accordingly. That assumption changes everything about how you write code, how you review code, and how you think about architecture.</p>
<p>“Where’s the timeout?” should be as automatic in code review as “where’s the null check?” “What happens when this dependency is down?” should be the first question in every architecture discussion, not the last. “What’s the fallback?” should be in every feature specification. “Why didn’t we fail gracefully?” should be a standard line item in every post-incident review.</p>
<blockquote>
<p>“Everyone has a plan until they get punched in the mouth.”* — Mike Tyson*</p>
</blockquote>
<p>Every service has a happy path. Resilience is what happens when the happy path gets punched in the mouth. And the time to plan for that punch is now — not when you’re lying on the canvas at 2 AM, watching your Grafana dashboard turn red one panel at a time.</p>
<p>Now, if you’ll excuse me, I need to go check the timeout configuration on that new service we deployed yesterday. I’m sure it’s fine. It’s always fine. Until it isn’t.</p>
]]></content:encoded>
  </item>
</channel>
</rss>
