“But, But… My Points!”: The Story Point Trap That’s Killing Your Engineering Culture
Or: Why Your Team Might Be Measuring Success While Missing the Point Entirely
You know that feeling when you realise something fundamental has gone wrong, but you can’t quite pinpoint when it happened? That moment when a well-intentioned practice has quietly mutated into something that actively works against you?
I had that moment a while ago. I was sitting with one of my engineers, discussing a small change to how we handle sprint carry-over. Nothing dramatic — just a suggestion that if a story isn’t 100% done, we shouldn’t count the “percentage complete” toward sprint completion. My reasoning was simple: we should optimise for done, not for half-done. Partial credit doesn’t incentivise finishing; it incentivises starting, and generally this makes for too much in progress work, and helps you ignore problems with flow.
His response stopped me cold.
“But, but… my points!”
Four words. That’s all it took to reveal that something had gone fundamentally sideways. Here was a talented engineer, genuinely distressed — not about delivering value to customers, not about solving business problems, but about a number on a dashboard. Story points had stopped being a planning tool and had become the purpose.
As the economist Charles Goodhart famously observed, “When a measure becomes a target, it ceases to be a good measure.” We’d done exactly that. We’d taken a rough estimation heuristic and transformed it into the scoreboard by which the team judged itself.
How We Got Here
The thing is, nobody wakes up and decides to optimise for meaningless metrics. It happens gradually, through a series of entirely reasonable decisions.
Product Owner needs visibility into team capacity? Start tracking velocity. Leadership wants to know if the team is “performing”? Compare velocity across sprints. Someone asks why velocity dipped? Well, now velocity is something you need to defend. And once you’re defending a number, you start optimising for the number.
The engineer I was talking to wasn’t being unreasonable. In his world, story points were success. The Product Owner measured his own success the right way — by whether his ideas achieved their intended business outcomes. But he measured the team’s success by sprint completion. Points delivered. Velocity maintained. The team existed to implement his ideas; he owned the results, they owned the throughput. Somewhere along the way, ownership of outcomes got completely separated from ownership of work. They’d become what Melissa Perri calls a “feature factory” — valued for their output, disconnected from the impact of what they produced.
The basketball coach John Wooden put it perfectly: “Never mistake activity for achievement.” We’d built an entire measurement system around activity and called it achievement.
The Output Trap
Here’s the uncomfortable truth: most engineering teams are measuring the wrong things. Not because they’re incompetent, but because outputs are easy to measure and outcomes are hard.
Consider what’s easy to count:
Story points completed per sprintNumber of deploymentsLines of code writtenFeatures shippedTickets closedNow consider what actually matters:
Did customer behaviour change?Did we reduce time-to-market for the business?Are engineers more productive than they were six months ago?Did we reduce incidents affecting partners?Are we actually solving the problem we set out to solve?The first list gives you a number by Friday. The second list requires you to wait, observe, and accept uncertainty. Guess which one ends up on the dashboard?
Peter Drucker warned us decades ago: “There is nothing so useless as doing efficiently that which should not be done at all.” Yet we’ve become remarkably efficient at measuring our efficiency at doing things that may or may not matter.
The Card Exercise
When I work with my managers on goal-setting — whether it’s OKRs, KPIs, or whatever flavour your organisation prefers — I’ve found that abstract discussions about “outcomes versus outputs” don’t land. People nod along, then go back to writing the same metrics they always have.
So I created a simple exercise. I give them a deck of cards, each with a different goal or measure written on it. Their job is to sort them into two piles: Outcome or Activity.
Try it yourself with these examples:
“Migrate all services to Kubernetes”
Outcome or activity?
That’s an activity. Classic “doing a thing” with no stated benefit. Why are we migrating? What improves when we’re done?
“Reduce deployment lead time from 2 weeks to 2 days”
That’s an outcome. Measurable, with clear value. We’ll know when we’ve achieved it, and we can tell if it’s working.
“Decouple the monolith”
This is where it gets interesting. Most people say outcome. It sounds like a goal. But it’s an activity dressed up in important-sounding language. The outcome would be something like “Enable the Supply and Booking teams to release independently” — that tells you why decoupling matters.
“Improve system observability”
Activity. Too vague to be an outcome. How would you know when you’ve achieved “improved observability”? What does success look like?
“Achieve loose coupling between services”
Still an activity, even though it sounds architectural and sophisticated. The question to ask: loose coupling enables what? Faster delivery? Fewer incidents? Those are the outcomes. That’s where the value is, this is what you should be measuring.
“Reduce mean time to recovery from 4 hours to 30 minutes”
Now we’re talking. That’s a measurable operational improvement with clear business value.
The pattern becomes obvious once you see it: activities describe what you’re doing; outcomes describe what changes because you did it.
The “So What?” Test
Here’s the simplest filter I’ve found: for any goal or metric, ask “so what?”
“We migrated to Kubernetes.” So what? “So we can… deploy faster?” So what? “So we can… get features to customers sooner?” So what? “So customer satisfaction improves and churn decreases.”
Now you’ve found the outcome. The migration was the activity. Customer retention is the outcome. Everything in between is either a stepping stone or a vanity metric.
The management thinker W. Edwards Deming spent his career arguing that “running a company on visible figures alone” leads to disaster. The figures we can easily see — story points, velocity, deployment counts — are often the least important. The figures that matter — customer behaviour change, business impact, developer satisfaction — require patience and interpretation.
What Actually Changed
Going back to my “but, but… my points!” moment: the real problem wasn’t the engineer. He was responding rationally to the incentive system we’d created. The problem was that we’d built a measurement culture that disconnected effort from impact.
Here’s what we changed:
We stopped reporting velocity upwards. Velocity is a team planning tool. The moment it becomes a performance metric, teams start gaming it — story point inflation, avoiding complex work, counting partial completion. Instead, we report on business outcomes: deployment frequency, incident rates, time to deliver customer-facing features.
We made carry-over painful. Not punitively, but visibly. Unfinished work stays on the board, staring at you, until it’s done done. No partial credit. This sounds harsh, but it actually freed teams to have honest conversations about scope and to push back on overcommitment. And watch out of this, too many things in progress at the same time, might mean its hiding queues that interrupt flow.
We focused team goals to customer outcomes. Every team has visibility into what happens after their code ships, we have an AB testing platform for this, everything is under a test, engineers own the results with Product, it’s one team. Did the partner feature reduce support tickets? Did the performance improvement change booking conversion? When engineers can see the impact of their work, story points become obviously absurd as a success measure.
We reframed what “success” means. Shipping a feature isn’t success. Shipping a feature that achieves its intended outcome is success. This subtle shift changes everything — from how you write acceptance criteria to how you decide when something is “done enough.”
The Danger of Activity Focus
There’s a deeper reason why outcome-thinking matters beyond just having better metrics. When you commit to an activity, you’re locked in. You’ve told everyone “we’re doing microservices” or “we’re migrating to the cloud,” and now your credibility is tied to completing that activity — even if you discover halfway through that it’s not working, that the assumptions were wrong, that there’s a better path.
When you commit to an outcome — “teams can deploy independently” or “we can recover from failures in under an hour” — you stay flexible. If Kubernetes doesn’t get you there, you can pivot. If the microservices approach is creating more problems than it solves, you can adjust. You’re measuring success by what improves, not by whether you completed the plan.
The general and strategist Helmuth von Moltke wrote that “no plan survives contact with the enemy.” In software, no architecture survives contact with production. Outcome-focus gives you permission to learn and adapt. Activity-focus locks you into the original plan regardless of what reality reveals.
A Question for Your Next Planning Session
Next time you’re setting goals for a quarter or reviewing OKRs, try this: for every item on the list, ask whether it passes the “so what?” test. Does it describe what you’re doing, or what changes because you did it?
“Implement feature flags across all services” — activity. “Reduce deployment risk so teams can ship daily with confidence” — outcome.
“Create an API gateway layer” — activity. “Enable partners to integrate with a single, stable interface” — outcome.
“Modernise our technology stack” — activity (and a vague one at that). “Reduce new engineer onboarding time from four weeks to one week” — outcome.
The words matter. The frame matters. Because what you measure is what you optimise for, and what you optimise for shapes what your team believes success looks like.
The Bottom Line
Story points aren’t evil. Velocity isn’t inherently misleading. Activities aren’t bad. But when we mistake the measure for the goal, when we optimise for the metric instead of the meaning behind it, we end up with engineers who are genuinely distressed about losing “their points” while the actual business outcomes go unexamined.
As the poet and philosopher Kahlil Gibran wrote — in an entirely different context but with perfect relevance — “We choose our joys and sorrows long before we experience them.” When we chose to measure story points, we were choosing, without realising it, to make points a source of joy and their absence a source of sorrow. We could have chosen to measure customer impact instead. We could have chosen outcomes.
The good news is, you can still choose. You can reframe what success means, reconnect effort to impact, and build a culture where “but, but… my points!” sounds as absurd as it should.
Now, if you’ll excuse me, I need to go remove a velocity chart from a dashboard. It’s been measuring our activity very precisely while telling us nothing about whether any of it mattered.