Beer & Servers Don't Mix

Silent Standby: The Invisible Tax on Your Engineering Velocity

Or: Why Your Sprint Is a Two-Week Hiding Spot for Waiting

It was 10:47 on a Wednesday morning, and Jeno had that look — the one where confidence hasn’t quite left the face but the jaw has tightened just enough that you know something’s wrong. He’d volunteered forty minutes earlier, during the planning session, with the easy assurance of someone who’d done this a hundred times: “I’ll knock out the backend change, should be quick, then I’ll help you guys with the front end work.” The whole team had one goal: finish this story by end of day. The TeamCity build had been running for fifteen minutes when it failed. He clicked Rerun. Went for a coffee. Came back thirty minutes later to a red build and a cold mug. Failed again. The room had that particular silence — not the productive kind, the kind where four engineers are watching the clock and pretending not to notice that their day’s goal is already slipping.

That silence wasn’t a bad build. It was the sound of an entire team’s momentum bleeding out through a flaky CI pipeline that nobody had fixed because nobody had ever had to.

As the physicist and management thinker Eliyahu Goldratt put it, “An hour lost at a bottleneck is an hour out of the entire system.” He was describing factory floors, but he might as well have been describing what happens when your engineer spends a morning fighting a build system instead of shipping the change your whole team is waiting on.

I want to name something that I don’t think has a name yet. I’ve been calling it Silent Standby — the invisible state where an engineer is technically “working” but actually waiting on someone or something. They don’t flag it as a blocker because it doesn’t feel dramatic enough. They context-switch to something else, lose focus, and the original task silently ages. Nobody tracks the cumulative cost. Nobody even notices it’s happening.

If my previous post on Code Entropy described how codebases silently degrade through small technical decisions, Silent Standby describes how organisations silently degrade through small invisible waits. Same mechanics, different substrate.

The Two-Week Anaesthetic

Here’s the thing about Jeno’s build problem: it wasn’t new. That CI pipeline had been flaky for months. But in a normal two-week sprint, the pattern looks completely different. You kick off the build, it fails, you rerun it, and while you wait you pick up another task. You check Slack. You review a PR. You answer a question. The build fails again — you rerun it, context-switch again. Over the course of two or three days, you get it through. Nobody notices. Nobody complains. The story gets marked as done within the sprint. Velocity looks healthy.

The sprint is an anaesthetic. It numbs you to the pain of waiting by giving you enough time to fill the gaps with other work. And because you’re always “busy,” the wait never surfaces as a problem.

We didn’t discover this by accident.

My former manager Pete had an exercise he’d run with teams — deceptively simple, almost casual. Take a medium-sized story, one that the team estimates at maybe two or three days of work, and try to finish it in a single day. Get the whole team in a room. Plan it in the morning. Ship it by end of day.

They almost always fail. But the failure is the point.

When we ran this for the first time, we knew the story required changes to two systems — a backend service and a front-end application. We had the backend system owner on our team that quarter, so code reviews were covered. Everything should have been straightforward.

Jeno took the backend change. Did the work, opened the PR, then went into TeamCity to run the build that would publish the NuGet package the BFF needed to consume. Fifteen minutes. Failed. Rerun. Coffee. Thirty minutes. Failed again. He repeated this pattern several times before he started investigating why it was flaky. Finally got it to work — then realised he’d missed a field in the contract, went back, made the change, triggered the build again. Failed. And again.

In a two-week sprint, this would have been invisible. Jeno would have kicked the build, gone to work on something else, checked back hours later, rerun it, moved on again. Over two or three days, it would have eventually worked. Nobody would have flagged it. Nobody would have tried to fix the underlying flakiness. It would have just been… how things are.

But compressed into a single day, with four other engineers sitting there unable to start the front-end work until that package published, the cost became impossible to ignore. The anaesthetic wore off.

What the Exercise Surfaces

Pete’s “One Story, One Day” exercise is really a diagnostic tool for Silent Standby. By compressing the timeline, every hidden wait becomes painfully visible. And the list is always longer than anyone expects:

The flaky CI pipeline that you just rerun a few times, and six hours later it’s ready. The code review ping-pong that stretches across six days because someone’s sick, or multiple teams need to approve a single change. The deployment you have to wait until the next day for because deployments only run daily. The multiple systems that need to merge and deploy in a specific order because your one story touches three separate codebases. The system that takes two to four days to get running on your laptop because you haven’t touched it in a month.

Each of these is a silent standby event. Individually, none of them feel like blockers. Collectively, they’re a tax on everything your team does.

You know the warning signs that a team is drowning in Silent Standby before you even run the exercise. At sprint start, each developer picks up a separate story and works in isolation. Stand-ups feel pointless because everyone’s working on different things and has nothing meaningful to coordinate. The team isn’t talking during the day — headphones on, everyone in their own world. And you hear a consistent refrain when you suggest collaboration: “We can’t get more than one person working on this.”

They’ve usually never tried.

The Science of the Switch

Here’s where it gets worse, because Silent Standby doesn’t just waste the time you spend waiting. It poisons the time after you stop waiting, too.

Dr. Sophie Leroy at the University of Washington has a term for this: attention residue. Her research shows that when you switch from Task A to Task B, part of your cognitive attention stays stuck on Task A. You’re not fully present on the new task. The more engaging or unfinished the original task, the stronger the residue.

Think about what this means for a developer who’s waiting on a build. They’ve got the entire architecture of their change loaded into working memory — the API contract, the downstream dependencies, the edge cases they still need to handle. The build fails. They can’t proceed. So they pick up a code review, or respond to a Slack thread, or start looking at a different story. But their brain is still partially running that first task in the background, like a browser tab you can’t close.

Research by Gloria Mark at UC Irvine found that it takes an average of twenty-three minutes to fully regain focus after an interruption. Gerald Weinberg’s data on context-switching in software teams suggests that when you assign a developer to two concurrent tasks, you get roughly 40% productivity on each — an immediate 20% loss just from the switching itself. Add a third task and you’re haemorrhaging capacity.

Now multiply this across every engineer on your team, every day, across every sprint. Those “minor” waits that everyone works around? They’re not free. They’re a compound tax on your team’s cognitive capacity, paid in bugs shipped, edge cases missed, and that general sense everyone has that “everything takes longer than it should.”

As Goldratt also observed, “Activating a resource and utilising a resource are not synonymous.” Your engineers look busy. Your velocity chart looks fine. But utilisation and productivity are not the same thing, and Silent Standby is the gap between them.

The 75-Day Merge Request

Let me give you a concrete example of what Silent Standby looks like at scale.

We had a Sourcegraph batch change that created a merge request in another team’s repository. A routine change — the kind that should take a reviewer thirty minutes. Nobody owned it. It sat there. For seventy-five days.

When someone finally noticed, the response from the owning team’s manager was illuminating: “Close it and create a new one.” Why? “It will impact our velocity metrics.”

Let me be clear about what happened here: a merge request was born from Silent Standby — no clear owner, no urgency, no one whose workflow it blocked enough to chase. It sat for two and a half months in a state of quiet neglect. And when it was finally discovered, the proposed solution was to hide the problem by resetting the metric. Close the old MR, open a new one, and the cycle time dashboard looks clean again.

I told the manager at the time: impacting your velocity metrics is not a reason to close an MR. They solved it by creating a new MR, which resets the metric and hides the problem. That’s not resolution — that’s gaming the number.

This is what happens when Silent Standby meets measurement culture. The wait becomes invisible, the metric becomes the goal, and the actual problem — that cross-team changes have no ownership model — never gets addressed.

Why You Don’t Notice

Silent Standby persists because of three reinforcing dynamics that conspire to keep it invisible.

The sprint buffer masks the cost. In a two-week sprint, a day of waiting gets absorbed. The work gets done eventually. The story moves to Done. Nobody asks whether it could have been done in a third of the time. The buffer exists precisely so that variability doesn’t hurt you — but it also means you never feel the variability, and therefore never fix the causes.

The bystander effect distributes responsibility. When a cross-team dependency creates a wait, it’s nobody’s emergency. The requesting team is busy with other work. The owning team has their own priorities. The MR sits in a queue that belongs to no one. In a co-located team, you’d walk over to someone’s desk and say, “Hey, can you review this? I’m blocked.” In a distributed, async-first world, you post a Slack message that gets buried under seventeen other notifications.

The “I’ll work on something else” reflex feels productive. This is the most insidious part. When you hit a wait, switching to another task feels like the responsible thing to do. You’re not sitting idle — you’re being efficient! Except you’re not. You’re paying the attention residue tax, fragmenting your focus, and ensuring that neither task gets your full cognitive bandwidth. But it feels productive, and feelings are what we optimise for when we don’t have data.

Making It Visible

You can’t fix what you can’t see. Here are three approaches we’ve used to drag Silent Standby into the light.

One Story, One Day

I’ve already described this, but I want to emphasise: the exercise isn’t about actually finishing the story in one day. It’s about what you learn when you try. The retro at the end answers one question: “What stopped us from going faster?” And the answers are always some variant of Silent Standby.

Run it once a quarter. Pick a real story, not a trivial one. Get the whole team in the room — physically, if possible. Plan it together in the morning. Watch what happens. The list of impediments you surface will fuel your improvement backlog for months.

MR Spread Analysis

Count the number of merge requests per feature or story, and look at how they spread across systems and teams. If a single story consistently requires PRs across three or four repos owned by different teams, you have a quantitative signal that your architecture is coupled in ways your org chart pretends it isn’t.

This isn’t a metric to optimise directly — you’re not trying to get the number to one. But it tells you where your service boundaries are wrong, where your teams are inadvertently creating cross-team dependencies, and where Silent Standby is most likely to accumulate.

The “One Story, One Day” exercise naturally surfaces this. When the team tries to finish a story in a single day and discovers they need changes in three systems owned by two other teams, the MR spread is the quantitative evidence of why they couldn’t finish.

Lead Time Decomposition

Stop measuring total lead time as a single number. Break it into active work time versus wait time. The ratio will shock you. In most teams I’ve worked with, the wait time dwarfs the work time by a factor of three to five. Your story takes ten days to deliver, but the actual hands-on-keyboard time is two days. The other eight? Silent Standby.

The Intra-Team Fix: Stop Being Strangers

Within a single team, Silent Standby is almost entirely solvable. The fix isn’t a tool or a process — it’s a shift in how people work together.

In my post on Mesh Programming, I argued that product engineering teams need to stop behaving like globally distributed open-source contributors and start behaving like people who sit next to each other. Define your contracts up front. Work on the same branch. Draw dependency lines between your tasks — literally, with post-it notes on a whiteboard — and identify where the serial chains are. Then break them.

The technique is simple: find the point in your dependency chain where you can insert a mock or a contract. Instead of waiting for the backend to be complete before starting the front end, create a controller that returns a mock view model. One hour of work, and suddenly your front-end and backend developers can work in parallel instead of serial. The task that would have taken thirty hours of elapsed time across a sequential chain can compress dramatically when you remove the waiting.

Go talk to the person. Make it synchronous. You are not a globally distributed team that works on the Linux kernel — you are a product engineering team. Act like one.

The Nuclear Option: Acceleration Week

If you want proof that eliminating Silent Standby works, run an Acceleration Week.

We did one — five days, twenty-five engineers from four companies, co-located in the same space. Same people, same tools, same technology they use every day. The only difference was proximity and focus.

The results were staggering. Engineers who normally took two or three days just to understand what needed to be done were getting feedback and direction in minutes. Problems that would have been Slack messages waiting for replies “maybe next week, because we’re busy” got solved on the spot. One engineer described it perfectly: “Less context switching. Normally, day to day, I’m doing like ten things at a time and nothing gets done.”

On the Monday of that week, we had a large incident caused by the exact architecture the group was there to replace. By Friday, they’d shipped the replacement. They also delivered a new deployment architecture, testing frameworks, open-source contributions, and upskilled fifteen engineers. In five days.

That’s what happens when you eliminate Silent Standby. The same people become dramatically more effective — not because they’re working harder, but because they’re not waiting.

It reminded me of what Agoda used to feel like when I first arrived. Everyone in our department was on level six, and we only took up half the floor, so you could see everyone from one end. If you were waiting for a code review, you wouldn’t use the PR bot to ping someone on Slack. You’d walk over to their desk and say, “Hey, can you do this code review for me? I’m waiting.” The wait evaporated because making it visible was effortless.

The Uncomfortable Question

Here’s the question I want to leave you with, and it’s not a comfortable one: what if the reason everything feels slow isn’t complexity, or technical debt, or insufficient tooling — but simply that your people are spending most of their time waiting for each other without anyone noticing?

As the coach John Wooden once said, “Never mistake activity for achievement.” Your team is active. Your Jira board is moving. Your velocity is stable. But underneath all that motion, how much of your engineering capacity is actually producing value, and how much is quietly haemorrhaging into the gaps between teams, between systems, between a failed build and a successful one?

Silent Standby won’t show up in your retrospectives because nobody thinks to mention it. It won’t show up in your metrics because you’re measuring the wrong things. It won’t show up in your sprint reviews because the work got done — eventually.

It shows up in the resignation letters, when engineers who are tired of feeling busy but unproductive decide to go somewhere they can actually ship things. It shows up in the slow, grinding deceleration that everyone notices but nobody can explain. It shows up in the gap between what your team could deliver and what they actually do.

Run the exercise. Compress a story into a day. Watch where the waiting lives. Then fix it.

Now, if you’ll excuse me, I need to go chase down a code review that’s been “almost done” for four days. But hey, at least my velocity looks great this sprint.