The Regenerate Button: Why Your Snapshot Tests Are Lying to You
Or: The Most Dangerous Command in Your Test Suite
You’ve run it. Don’t pretend you haven’t. That magical command that makes all those failing snapshot tests go away. --updateSnapshot. --update-snapshots. Whatever flavour your framework prefers. You've run it, committed the changes, and moved on with your day. We all have.
And that’s exactly the problem.
As the legendary basketball coach John Wooden once observed, “The true test of a man’s character is what he does when no one is watching.” In our world, the true test of a test suite’s value is what happens when those tests fail and nobody’s looking over your shoulder. Spoiler: we regenerate.
The Conversation That Changed How I Think About Tests
Years ago, when we’d just started using React, I had a conversation with one of my developers — let’s call him Toey — about snapshot tests. We’d been having debates about their usefulness, the kind of debates that happen in every engineering team that adopts a new testing pattern.
So I asked him directly: “Toey, what do you do when the snapshot tests fail?”
He looked at me with complete sincerity and said, “Joel, I have this command someone showed me.” And he ran that command, and the tests passed again.
The command, of course, was to regenerate the snapshots.
I didn’t need to say anything else. He’d just proven my point about how useless they were in practice, regardless of their theoretical value.
The Theory vs. The Reality
Now, before the snapshot testing advocates come for me with pitchforks, let me be clear: I’ve had many talented engineers explain the value proposition to me. And they’re not wrong. In theory, snapshot tests are elegant. They capture the output of a component at a point in time, and any deviation from that snapshot signals a change that needs human review. It’s regression testing with minimal effort.
The theory is beautiful. The reality is that regenerate command.
The economist Thomas Sowell put it perfectly: “There are no solutions, only trade-offs.” The trade-off with snapshot tests is this: they require discipline, and discipline is exactly what modern product engineering culture actively works against.
Think about your average sprint. Your manager is asking for status updates. Your PO is wondering why this feature is taking so long. You’ve got three other PRs waiting for review. And suddenly your CI pipeline goes red because seventeen snapshot tests failed after you changed the padding on a button.
What do you do? Do you carefully review each snapshot diff, considering the implications of every changed pixel? Or do you run the regenerate command, eyeball the changes in the PR diff, and move on?
Be honest.
The Scale Problem
Here’s what the snapshot testing evangelists often miss: discipline doesn’t scale.
When you’re a small team — five engineers, maybe ten — you can maintain the discipline. You know every component intimately. You review snapshot changes carefully because you wrote most of them. The context is always fresh.
But when your repository is processing its 120th merge request this week from the 100th developer? When half the team has never seen the component that just failed? When the person who wrote that snapshot left the company two years ago?
The regenerate command isn’t a failure of character. It’s a rational response to an impossible situation. You cannot expect humans to carefully review changes to files they don’t understand, in systems they didn’t build, under deadline pressure that never relents.
This is entropy in action. Tests that require extra effort will eventually fall to it, especially in larger organisations with significant cross-system work. Every time someone runs that regenerate command without truly reviewing the changes — and it doesn’t happen every day, but it happens enough — the tests lose a little more of their value. Eventually, they’re not testing anything meaningful. They’re just generating green checkmarks.
Meet the New Boss, Same as the Old Boss
I’m thinking about all this today because we’re living through the same cycle again, just with different technology.
We’re not using React snapshots anymore. The new kid on the block is Playwright component tests with screenshot comparisons. Same concept, shinier packaging. Instead of comparing serialised component trees, we’re comparing actual rendered pixels.
In some ways, it’s better. You’re testing what users actually see, not an abstraction of what they might see. Visual regression testing catches things that DOM snapshots miss.
In other ways, it’s worse. Much worse.
Screenshots are environment-sensitive in ways that DOM snapshots aren’t. Linux renders fonts differently than Mac, which renders differently than Windows. A screenshot taken on one developer’s machine won’t match one taken on another’s, even with identical code. Rendering engines have subtle differences. Anti-aliasing varies. That pixel-perfect comparison you’re running? It’s comparing apples to slightly different apples.
And the regenerate problem? It’s not just the same — it’s amplified.
Time Bombs in the Test Suite
Just last week, I discovered something that perfectly illustrates this problem. We had tests with hardcoded dates — 2025-1-15 and 2025-1-20 — embedded in test selectors. You know what those are? Time bombs. The moment we passed January 2025, those tests started failing because the calendar components wouldn't render dates that were now in the past.
The fix is straightforward: Playwright has a clock API that lets you mock the system date, so tests always run in a controlled time context. Problem solved, lesson learned.
But here’s the kicker: I needed to bump the Playwright version to access that clock API properly. And what happened when I bumped the version?
All the screenshot tests broke.
Every. Single. One.
Not because the code changed. Not because there was a bug. Because the rendering engine produced slightly different output at the pixel level. The tests were doing exactly what they were designed to do — detecting visual changes — and yet the “changes” they detected were completely meaningless.
Guess what command I ran next?
The Deeper Problem
The issue isn’t snapshot tests or screenshot tests specifically. The issue is any testing approach that requires consistent human judgment under inconsistent conditions.
The philosopher Blaise Pascal wrote, “All of humanity’s problems stem from man’s inability to sit quietly in a room alone.” Our testing problems stem from our inability to sit quietly with a failing test and truly understand whether the failure matters.
We want tests that run automatically, fail meaningfully, and require minimal intervention. Snapshot tests promise the first two but absolutely require the third. And that requirement is incompatible with how modern software development actually works.
What Actually Works
I’m not saying visual regression testing is worthless. I’m saying we need to be honest about what it can and can’t do.
Here’s what I’ve seen work:
Component-level visual tests with tight scope. Test individual components in isolation, not entire pages. Smaller surfaces mean fewer spurious failures and more meaningful diffs when things do change.
Deterministic environments. Run your screenshot tests in containers with identical configurations. Eliminate the Linux/Mac/Windows problem by eliminating the variation. This is table stakes, and it’s surprising how many teams skip it.
Semantic assertions over pixel comparisons where possible. Does the button exist? Is the text correct? Is the layout flowing in the right direction? These assertions don’t break when someone updates a dependency.
Treat regeneration as a code smell, not a solution. If you’re regenerating snapshots frequently, that’s a signal that your tests aren’t providing value proportional to their maintenance cost. Either fix the underlying flakiness or delete the tests.
The Uncomfortable Truth
We want testing to be solved. We want to write tests once, have them catch bugs forever, and never think about them again. Snapshot testing feels like it offers this — set it and forget it, let the machine do the work.
But the machine doesn’t do the work. The machine generates failures. Humans have to do the work of interpreting those failures, and humans under pressure take shortcuts. This isn’t a moral failing; it’s a design failure. We’ve built processes that require superhuman discipline and then act surprised when humans can’t maintain it.
If that commitment doesn’t exist — and in most organisations, for most tests, it doesn’t — then the tests aren’t providing the value you think they are. They’re just making your CI pipeline take longer.
The Bottom Line
Snapshot tests are a good idea that fails in practice because they require a level of discipline that modern engineering organisations don’t support. Screenshot tests inherit all the same problems and add new ones. The regenerate command isn’t a bug in developer behaviour; it’s a rational response to a testing approach that demands too much and provides too little.
This doesn’t mean you should abandon visual regression testing entirely. It means you should be honest about its limitations and design your testing strategy accordingly. Use it sparingly, in controlled environments, for components where visual accuracy genuinely matters. And when you find yourself reaching for that regenerate command, ask yourself: is this test actually providing value, or is it just generating work?
The answer might be uncomfortable. But it’s better than lying to yourself with green checkmarks that don’t mean anything.
Now, if you’ll excuse me, I need to go regenerate some screenshots. I bumped a Playwright version and, well, you know how it goes.