Your Performance Dashboard Is Lying to You
Or: How to Actually See the Problem Before You Try to Fix It
The numbers were green. They had been green for three weeks. Somchai had pulled up the dashboard at 10:23 on a Thursday morning — Thai iced tea sweating condensation onto the vinyl floor beside his desk, the air conditioning on level 6 doing its usual indifferent job — and everything looked fine. Load time: fine. Error rate: fine. The graph lines sat flat and well-behaved, well below their thresholds. He’d closed the laptop a few minutes later and headed to planning. The user complaints had been sitting in the feedback queue the whole time.
This is the trap. Not that you’re not measuring. You probably are. The trap is that your dashboard is measuring the right things in the wrong way, or the wrong things entirely, and the gap between “the metrics are good” and “users are happy” is where performance work goes to die.
Before you can fix a performance problem, you have to be able to see it. That sounds obvious until you’ve spent three days optimising something your users don’t experience, or celebrated a P50 improvement while your slowest users got slower. This post is about the instruments — what they measure, what they miss, and which number from your data actually tells you something useful.
Two Kinds of Measurement, Two Different Questions
There are two fundamentally different approaches to measuring web performance, and confusing them is the source of a lot of wasted effort.
Real User Monitoring (RUM) captures what actually happens when real people load your pages. Real devices, real networks, real geographic distribution, real everything. It’s the gold standard for answering “what is my users’ actual experience right now?” The limitation is that it’s retrospective — you only know about a problem after users have already experienced it. And it only works on pages you control.
Synthetic monitoring runs automated browsers in controlled conditions — a scripted agent loads your page, measures it, and reports. It catches regressions before users see them, produces repeatable results you can track over time, and can be pointed at any URL. That last part is important: because synthetic monitoring is just “open a browser and load a page,” you can run it against your competitors just as easily as against yourself.
The honest summary: for measuring production performance, RUM is your primary instrument. Synthetic is how you catch regressions before they ship and how you know where you stand relative to the market. They’re complementary, not competing.
One important caveat: if your team runs A/B experiments — and at any serious product company, you should be — synthetic monitoring in CI becomes significantly less useful as a regression gate. Which variant do you test? The control? The treatment? Both? If a performance regression is intentional in one variant to test a hypothesis, your pipeline breaks. If you only test the control, you’re blind to what you’re actually shipping to half your users. In practice, synthetic in CI is almost useless for performance regression detection in an experimentation-heavy environment. To do it properly, your pipeline would need to know which feature flags a given PR affects, spin up synthetic runs for both the control and treatment configurations, and compare them — for every PR that touches an experiment. Almost no team does this, because almost no team has built the tooling to make it tractable. What you end up with instead is a synthetic run against your default configuration, which may bear little resemblance to what half your users are actually loading. RUM on your experiment cohorts is the only instrument that tells you whether a given variant is actually fast for real users. Don’t let a green CI pipeline give you false confidence when the thing users are experiencing is a variant your synthetic agent never loaded.
The Metric Question: What Are You Actually Measuring?
Here is where a lot of teams are quietly in trouble — not because they aren’t measuring, but because they’re measuring something their users don’t experience.
Performance metrics have a history, and it matters because the metric your team uses today is probably a product of when someone last revisited the question rather than a deliberate current decision.
The browser event era
For years, the standard metrics were DOMContentLoaded and load — browser events that fire as the document is parsed and resources are fetched. These made perfect sense for server-rendered pages where the content was in the initial HTML. If your page loaded, the HTML had your content.
Then React happened. And Angular. And Vue. And the SPA pattern spread across the industry. On a single-page application, the initial HTML is a shell. The content renders afterwards, in JavaScript, after the framework initialises. DOMContentLoaded fires on an empty page. load fires on an empty page. The metrics stay green. Users see a blank screen.
We hit this during our React migration. The dashboard showed no degradation. The pages were slower. Both things were simultaneously true because we were measuring the wrong moment entirely.
Core Web Vitals
In 2020, Google shipped Core Web Vitals: LCP (Largest Contentful Paint), CLS (Cumulative Layout Shift), and INP (Interaction to Next Paint, which replaced FID in 2024). These metrics are backed by large-scale research on what actually correlates with how satisfied users are with a page. They’re tied to search ranking. P75 is the standard reporting threshold. For consumer web products — ecommerce, content sites, booking flows — they’re the right starting point.
LCP measures when the largest element on screen has painted. For a consumer page that’s usually the product image, the hero content — the thing the user came to see. LCP is a good proxy for “is this page useful yet?” when you’re building browsing experiences.
CLS measures how much the layout shifts unexpectedly after it first appears. The number that tells you whether the page jumps around as images and ads load and knock your content out from under the user’s cursor.
INP measures how quickly the browser shows a visible response to user input — and we’ll come back to this one in detail because it catches a pattern that most teams have in their codebase right now and don’t know about.
Where Core Web Vitals don’t tell the whole story
LCP measures the largest painted element. For a consumer page, that’s the product image. For a B2B data application — a pricing tool, an inventory manager, an operations queue — the largest element might be a page header or a skeleton loader. LCP fires. The data table the user actually needs hasn’t rendered yet. The user is staring at a chrome frame with no useful information. The metric is green.
This isn’t a flaw in Core Web Vitals. It’s a context mismatch. CWV was designed for consumer ecommerce. Applied to data-heavy B2B interfaces without adjustment, it measures the wrong moment.
The solution is to instrument the page yourself. We built a component — a Higher Order Component wrapper — where developers nominate the element that defines “this page is ready to use.” One wrap around the primary data table, and an event fires when it mounts. The metric captures “data the user needs is available” rather than “the largest element has painted,” which for a loading skeleton are two very different things.
The principle is simple: standard metrics are right for the context they were designed for. When your context differs, you extend, you don’t just accept what doesn’t fit.
As the statistician George Box famously observed: “All models are wrong, but some are useful.” The same is true of metrics. Use the useful ones. Know where they stop being useful.
The Percentile Debate: Which Number From Your Data Actually Matters?
You have your RUM data. Thousands of real user sessions, a latency distribution. Which number do you put on the dashboard?
This question has a right answer, a wrong answer, and a historical answer that keeps changing — which is itself the insight.
The wrong answer: average latency. An average is a number no real user experiences. It’s distorted by outliers in both directions — a large volume of very fast requests can make a painful 95th percentile invisible, and a handful of extreme outliers can make a healthy typical experience look bad. If your performance dashboard is reporting average load time, you are flying with a broken altimeter. Percentiles are the only meaningful representation of a latency distribution.
The physicist Richard Feynman had a principle he applied to physics problems that applies here: “The first principle is that you must not fool yourself — and you are the easiest person to fool.” Averages fool you. Percentiles don’t.
What each percentile actually tells you:
Percentile What it shows The risk P50 (median) The typical user’s experience Masks the tail completely P75 Google’s CWV standard May miss genuine pain for a significant minority P90 Covers nearly everyone engineering can help Historically argued as “too noisy” P95 Near-universal experience Five percent exclusion is still large at scale P99 Right for internal service-to-service calls Dominated by noise on the public internet
The key insight: the right percentile is not fixed. It’s a function of how much of your tail reflects engineering problems versus environmental noise. In markets with inconsistent mobile infrastructure, the P90 tail is heavily influenced by conditions no engineering team can fix — report P90 and you’re optimising for the weather. In markets where networks have matured and device capability has standardised, the same tail is dominated by rendering and JavaScript execution problems that engineering absolutely can fix — and P75 may be letting you off too easily.
The question isn’t “which percentile should we use?” It’s “at what percentile does the tail stop being signal and start being noise?”
A practical two-tier model:
For consumer web RUM, P90 is a reasonable target if your markets are reasonably well-connected. You’re reaching nearly everyone engineering can help, without letting uncontrollable outliers drive the roadmap.
For internal service-to-service calls, P99 is right. You control both ends of the call. There’s no meaningful network variance between your own services. At P99, every slow request is an engineering problem, not a network one.
INP: The Metric That Caught the Loading Screen Trick
INP deserves its own section because it catches something common that the previous generation of metrics missed entirely.
FID, the metric INP replaced in March 2024, measured how quickly the browser started handling an input — specifically, the very first user interaction on a page. INP measures how quickly the browser showed a visible response to the input. And it measures all interactions throughout the visit, not just the first one.
The practical difference: a page that swallows a click, starts an async operation, and shows a spinner 500ms later has perfect FID and terrible INP. The old metric couldn’t see this pattern. The new one can.
Good INP is under 200ms at P75. Needs improvement: 200–500ms. Poor: over 500ms.
The implementation detail most teams get wrong is the sequence. A loading indicator must appear before async work begins:
- User clicks- Button immediately shows loading state — this is the “next paint” INP measures- Async work begins If your async operation starts before you update the DOM, INP measures the full async duration. A 500ms API call with a spinner that appears after it completes is 500ms INP, not “near-instant with a spinner.” The user clicked, saw nothing, and already clicked again.
This matters beyond UX. When users don’t see a response to their input, they re-click. They interpret a non-responsive button as “my click didn’t register.” At scale, re-clicks are not just a user experience problem — they’re a reliability problem. Multiplied input events hitting your backend at unpredictable rates, particularly during degraded performance when the response takes even longer than usual, when users are most likely to re-click.
Consider a search page with mediocre INP — no loading indicators, no button state change on click. It was a known issue that hadn’t risen to priority. A query plan regression caused the underlying database query to run slower. The page got slower. With no visual feedback on click — no spinner, no disabled state, nothing — users had no way to know whether their click had registered at all. So they clicked again. And again. Each user convinced their first click had silently failed, firing the same expensive search request two, three, four times. Re-clicks multiplied the load on an already-degraded query. What would have been a slow-page incident became an outage. The root cause was the query regression. The thing that turned degraded into down was the pre-existing INP problem that had been sitting in the backlog.
Fix the loading indicator before it becomes the thing that turns your next slow incident into your next outage.
Where This Leaves You
Before you open a performance ticket, before you start profiling, before you schedule a sprint around “performance improvements,” ask three questions:
Are you measuring what users actually experience? If you’re on a React SPA and your primary metric is DOMContentLoaded, the answer is no. If your B2B application's "ready" moment is when the data table mounts and you're measuring LCP of a skeleton loader, the answer is no.
Are you reporting the right percentile? If your dashboard shows averages, rebuild it. If it shows P50, you’re measuring your median user while your slowest users — who are more likely to churn, more likely to complain, and sometimes more likely to be your highest-value segments in B2B — are invisible.
Do you know what your competitors look like on the same metrics? Not to obsess over them, but because there is no absolute threshold at which performance “passes.” There’s only faster or slower than the alternatives your users have. Synthetic monitoring against a few benchmark URLs costs almost nothing and tells you whether your users are tolerating you or choosing you.
As Einstein put it — with characteristic economy — “The important thing is to not stop questioning.” Your monitoring stack is only as good as your willingness to interrogate it. RUM, synthetic, the right metric for the right surface, the right percentile for the right context. None of them works alone. Together, they tell you what’s actually happening — but only if you keep asking whether what you’re seeing is the whole picture.
The dashboard that says everything is fine while users are complaining isn’t a mystery. It’s a measurement problem. And measurement problems are the most fixable kind — once you can see what’s wrong, you can fix it. The hard part is admitting that the green numbers might not mean what you think they mean.
In the next post, we’ll look at where performance problems actually hide in a frontend codebase — because once you can see the metrics clearly, the next question is what to do about them.
This is the first in a three-part series on web performance engineering. Post 2 covers where performance problems actually live in the frontend. Post 3 covers how to build a performance culture that sticks.