Beer & Servers Don't Mix

Where Web Performance Actually Lives

Or: Why Your Frontend Engineer Can’t Fix What Isn’t a Frontend Problem

Namfon had done everything right. The Lighthouse score sat at 98. Images were lazy-loaded, the critical CSS was inlined, render-blocking scripts had been hunted down and eliminated. The lab numbers were clean. But the field data — real users, real devices, real networks — told a different story, and it had been telling it for weeks.

It was around 11am on a Wednesday, the kind of Bangkok morning where the AC on level 6 was already losing the argument with the sun coming through the floor-to-ceiling glass. The open-plan floor hummed — standups bleeding into each other from three directions, someone’s Teams notification going off every forty seconds, the faint plastic smell of fresh Thai tea from the kitchen drifting past. Namfon’s desk had that specific archaeology of someone who hadn’t quite left a problem: a cold cup of something she’d forgotten to drink, three browser tabs frozen mid-audit, a sticky note with “WHY” written on it and nothing else. She leaned back from her monitors, crossed her arms, and said to no one in particular, “I don’t know where else to look.”

She wasn’t wrong. She just wasn’t looking in the right place — because the right place wasn’t in the browser at all.

This is the second post in a three-part series on web performance. Post 1 covered how to see the problem — the metrics, the instrumentation, the specific incident that made me take INP seriously. This post is about where the fix actually lives once you can see it.

The short answer: usually not in the browser.

The browser is the surface where milliseconds are felt. It’s where Lighthouse runs, where your product manager takes screenshots of loading spinners, where the user’s frustration lives. But in a modern microservices architecture, the browser is often just the messenger — faithfully reporting the outcome of decisions made much further upstream. Decisions about how your BFF aggregates data, how your bundles are structured, and whether your teams own performance or just blame each other for it.

Let’s go through where the leverage actually is.

The BFF Layer: Sequential vs Parallel Is the Highest-Leverage Intervention You’re Probably Not Making

A rich page in a microservices architecture needs data from multiple services. User preferences, pricing, availability, promotional content — each comes from a different place. The Backend for Frontend (BFF) pattern exists to aggregate these into a single, page-optimised response, so the browser makes one clean API call rather than juggling five or ten.

But here’s the thing about that BFF call: the way it’s written internally determines whether your page loads in 100ms or 500ms, and it usually has nothing to do with any of the individual services.

If the BFF fetches those services sequentially — await serviceA(), then await serviceB(), then await serviceC() — the page response time is the sum of all service latencies. Five services at 100ms each: 500ms. If it fetches them in parallel, the response time is the maximum of all service latencies. Five services at 100ms each: 100ms.

Same services. Same backend. Same frontend. Five times faster.

The fix is small code: Promise.all([serviceA(), serviceB(), serviceC()]) in JavaScript, Task.WhenAll() in C#. The reason it doesn't get written this way the first time is simple — sequential awaits are easier to read and reason about. You write what makes sense in the moment, and the performance implication is invisible until you have real load and real data.

The real BFF, of course, is more nuanced than “just parallelise everything.” Some calls have dependencies: service B might genuinely need the result of service A before it can run. The right mental model isn’t “full sequential” or “full parallel” — it’s a dependency graph. What can run independently of what? Group the independent calls together, chain only where you actually have to. That analysis — mapping your BFF call graph and identifying what’s blocking unnecessarily — is one of the most underappreciated architectural tasks in frontend engineering, and it’s almost always done either too late or not at all.

One more thing that often gets missed before we even get to the BFF: browsers enforce a hard limit of six parallel HTTP/1.1 connections per domain. If your API calls and your static assets are competing for the same six slots, they’re throttling each other at the browser level. Serving your static assets — JS bundles, CSS, images — from a dedicated CDN domain frees those connection slots entirely for API traffic. This looks like a deployment detail. It’s actually a browser constraint optimisation, and it’s the reason you want to solve this at the BFF before you even think about making the frontend juggle multiple service calls directly.

Our BFFs at Agoda are .NET (C#) and Kotlin — not for any philosophical reason, but because both runtimes handle concurrent aggregation work at scale more efficiently than Node’s single-threaded event loop. When your BFF is doing significant parallel I/O and aggregation simultaneously, the runtime choice matters. Sam Newman’s original BFF pattern write-up calls out parallel orchestration as a core concern of the pattern, not an optimisation to layer on later. We’ve found that’s correct.

The Loading State You’re Getting Wrong (And Why It Matters at Scale)

This one is small, specific, and embarrassingly common.

INP — Interaction to Next Paint — measures the time between a user doing something and the browser visually acknowledging it. The key word is “visually.” INP doesn’t measure how long your async work takes. It measures how quickly the next paint after interaction happens.

The correct sequence when a user clicks a button:

  • User clicks- Button immediately shows a loading state — this is the next paint INP measures- Async work begins The common (wrong) sequence:

  • User clicks- Async work begins- Button shows loading state after the first state update resolves If async work starts before the DOM update, INP measures the entire duration from click to visible response — which includes the async work itself. A loading spinner that appears 300ms after click because it’s set inside a .then() is 300ms INP, not near-instant feedback. The fix is a state update before the async call, not inside it.

This matters beyond user experience scores. At significant traffic volumes, a page that doesn’t visibly acknowledge clicks creates re-click behaviour. Users interpret a non-responsive button as “my click didn’t register” and click again. Each re-click is another request. On expensive operations — search queries, booking submissions — a mediocre INP score can turn a slow-but-functional page into an overloaded-and-down page under the right load conditions. The web.dev INP documentation is the right reference here, and it’s worth actually reading rather than just targeting the number.

The DOM Nobody’s Looking At

Here’s one that comes up repeatedly in INP investigations and almost never gets caught in a Lighthouse run: engineers rendering DOM that users can’t see.

It sounds absurd when you say it plainly. Why would you render a thousand rows when the user can only see twenty? But it happens constantly, and it’s almost always unintentional. A search results page renders the full list. A feed renders every item. A data table renders all rows up front because the original dataset was small and nobody noticed when it grew. The page looks fine. Scrolling feels fine. And then, at some point, React needs to reconcile a state change — a filter selection, a price update, a hover interaction — and it dutifully re-renders every node in that enormous DOM. Including the eight hundred rows the user can’t see because they’re below the fold.

INP spikes. Engineers look at the interaction code and find nothing obviously wrong. The component is clean. The logic is simple. The problem isn’t the code — it’s the surface area the code is operating on.

DOM size is a multiplier on everything. The larger the DOM, the more expensive every rerender, every style recalculation, every layout pass. Browsers aren’t infinitely fast at traversing and updating nodes, and React’s reconciliation, efficient as it is, still has to do work proportional to what’s mounted. Mounting less is always faster than reconciling faster.

The solution that actually works at scale is virtualisation — only rendering what’s visible in the viewport and dynamically creating and destroying nodes as the user scrolls. TanStack Virtual is the library worth knowing here: it handles the windowing logic without prescribing your rendering approach, works with any framework, and manages the complexity of variable-height items that makes naive virtualisation implementations fall apart.

The pattern is straightforward in principle: you maintain a full list in memory, but at any given moment only a small window of that list exists in the DOM. As the user scrolls, nodes at the top get destroyed, nodes at the bottom get created. The user never notices. React rerenders become fast again because there are thirty nodes to reconcile instead of three thousand.

It won’t fix every INP problem. But if you have a slow interaction on a page with large lists or tables and you haven’t looked at DOM size yet, look there first. You might find the culprit isn’t the interaction at all.

As the statistician W. Edwards Deming observed, “Every system is perfectly designed to get the results it gets.” A high-tempo A/B testing culture is perfectly designed to accumulate JavaScript bundle bloat. Not because anyone is being careless — but because of how the system works.

Every A/B test requires both variants to be present in the bundle. The server doesn’t know at build time which variant a given user will see, so both ship. One test: a small amount of extra code. Twenty concurrent tests: a meaningful amount of dead code delivered to every user on every page load, where “dead” means “code the current user will never execute.”

This is the structural cost of experimentation speed. It’s not a mistake. But it compounds. Experiments conclude, their code should be removed, and in practice — honestly, in most organisations — it doesn’t get cleaned up immediately. The cleanup ticket sits in the backlog. The next sprint has higher priorities. The experiment code becomes part of the background noise, and then becomes load-bearing architecture because everyone’s forgotten what it was.

Three approaches exist for managing this:

Discipline: Remove experiment code when tests conclude. Assign ownership. Make it a deployment blocker. This works at small scale. It degrades at high experiment velocity because the process burden scales with the number of experiments, not the team size.

Compression Dictionary Transport (CDT): A relatively new HTTP standard, supported in Chrome and Edge from version 130, that uses a previously-cached version of a resource as the compression dictionary for the new version. Instead of downloading the full updated bundle on each deployment, the browser receives a compressed diff against what it already has. This doesn’t reduce what’s in the bundle, but it dramatically reduces what needs to be transmitted on update. Worth noting: CDT is most effective when you have regular returning users — B2B applications, social platforms, tools where daily or weekly users are common. For ecommerce, where a significant portion of users arrive for the first time or after long gaps, the benefit is reduced. The Chrome developers overview is the right starting point.

Edge tree shaking per user: Knowing a user’s A/B experiment assignments at request time — which an edge worker can do — you can strip the dead-variant code before serving the bundle. Each user receives only the code their session will execute. This is the most aggressive solution and the most complex to implement. Zephyr Cloud implements this as a bundler plugin plus edge worker and is worth examining if bundle size is a genuine bottleneck for you.

The honest position: most teams are running on discipline, watching it slowly fail, and haven’t yet seriously evaluated the alternatives.

The GraphQL Upstream Problem

We’re conditioned to optimise response payload size. Smaller responses download faster — this is obvious, it’s correct, and it’s incomplete.

HTTP connections are asymmetric. Significantly more downstream bandwidth is available than upstream. On mobile, particularly on inconsistent networks, upstream is a real constraint. In GraphQL-heavy architectures, the query body going up can itself be several kilobytes of JSON per request. On desktop, on a good connection, this is invisible. On mobile, at scale, it isn’t.

Our mobile teams at Agoda address this with request body compression on GraphQL queries — roughly 7:1 compression ratios on upstream payloads. Native mobile apps give you direct control over request construction, which makes this reasonably straightforward to implement. Web browsers don’t support request body compression the same way without explicit server negotiation, which makes this a mobile-first optimisation rather than a universal one.

The principle is worth holding onto regardless: in request/response cycles, both directions have cost. Your instinct to optimise response size is right. It’s just not the only thing worth optimising.

Performance Is a Culture, Not a Team

Here’s the failure mode that nobody wants to name because it feels unfair to the people working hard inside it.

Performance becomes a problem. Leadership creates a performance team to fix it. The performance team works hard, ships improvements, publishes dashboards. Meanwhile, every other engineering team continues making decisions — architectural choices, new features, third-party integrations, BFF call patterns — without performance as a first-class consideration. The performance team spends most of its time catching up with regressions introduced by everyone else.

Then the resentment sets in, from both sides. The performance team feels like they’re running up a down escalator. The other teams develop the attitude that kills this whole model: “Don’t we have a team for that?”

As the surgeon and author Atul Gawande wrote about failure in complex systems: “The problem is rarely a lack of effort. The problem is that the system doesn’t make the right behaviour the easy behaviour.” Creating a dedicated performance team doesn’t make good performance decisions the easy behaviour for everyone else. It makes outsourcing those decisions the easy behaviour — which is the opposite of what you need.

Performance in a large engineering organisation is downstream of every decision every other team makes. Schema choices, component render patterns, BFF aggregation structure, experiment accumulation, API payload design — none of these are owned by a performance team. All of them affect performance. A team that fixes performance problems is a crutch. A team that enables other teams to fix their own performance problems is a multiplier.

The model that works looks like this: the performance team sets and maintains SLAs with product teams, so performance thresholds are commitments rather than aspirations. They build and maintain the tooling — RUM instrumentation, dashboards, custom metrics — so teams can see their own performance without needing the performance team’s involvement. They sit in on architecture reviews early, when the cost of changing a bad decision is low, not after the regression has shipped. They run training — not on Lighthouse scores, but on how browsers actually work, why the loading-state sequence matters, what sequential BFF calls are costing.

The goal is teams that know how to make their pages fast and keep them that way. Not faster pages that drift back to slow once the performance team moves on.

What This Means in Practice

Pulling back to the original scene: Namfon’s Lighthouse score was 98. The page was slow. Both things were true simultaneously, and understanding why requires holding two ideas at once.

The browser is where performance is experienced. It is not where performance is determined.

The BFF’s aggregation pattern determines whether you pay 100ms or 500ms before the browser has anything to work with. The bundle’s structure determines how much code a first-time visitor has to parse. The loading state sequence determines whether INP is measured in milliseconds or seconds. None of these show up as “frontend problems” on Lighthouse.

The fix Namfon needed wasn’t a better Lighthouse audit. It was a conversation with the team that owned the BFF — a look at the call graph, a question about what was running sequentially that didn’t need to be.

That conversation is harder than running a Lighthouse report. It crosses team boundaries. It requires shared ownership of a metric that nobody fully owns.

That’s precisely why it’s worth having.

Post 3 of this series covers how you know a fix is actually working — measurement, baselines, and the specific metrics that tell you whether you’re improving or just making a different kind of slow.