Beer & Servers Don't Mix

Designing APIs That Outlive Your Current Sprint

Or: Why Your API Contract Is the Most Important Code You’ll Never Unit Test

It was 11:47 on a Tuesday morning when the war room on level 7 filled up. The Grafana dashboard on the NOC wall had gone from a lazy green river to an angry red waterfall — a massive spike in 4xx responses from the BFF, users hitting errors across multiple pages. Someone had already grabbed the Thai iced teas they’d never finish. The vinyl flooring squeaked under rolling chairs being pulled in too fast. Within ten minutes, we’d traced it to a BFF deployment: an intern had renamed a field from cust_id to customer_id. Reasonable change. Good naming, even. Their tests passed. Their code review passed. But every browser out there still running the previous version of the SPA — which was all of them — was parsing a field that no longer existed.

That’s not a story about an intern making a mistake. That’s a story about an organisation that treated an API like an implementation detail instead of what it actually is: a promise to every browser tab still open from this morning.

The Promise Nobody Reads

As the investor Warren Buffett once observed, “It takes 20 years to build a reputation and five minutes to ruin it.” He was talking about business, but he might as well have been describing your API contract. You spend months building trust with your consumers — they wire up their parsing, their error handling, their retry logic — and then one renamed field on a Tuesday morning means every user whose browser hasn’t fetched the latest JavaScript is now staring at a broken page.

Here’s the thing that took me years to internalise: your API is not your implementation. Your implementation is the messy, evolving, refactorable stuff behind the curtain. Your API is the curtain itself — the surface that every other team builds against, the contract that makes independent deployability possible. The moment you break that contract, you’re not deploying independently anymore. You’re coordinating. And coordination, as anyone who’s tried to get six teams to deploy in sequence on a Friday afternoon knows, is where velocity goes to die.

This post isn’t an API design textbook. It’s a set of practical checks — things you can run through in a code review in under ten minutes — that separate APIs that survive contact with the real world from APIs that generate a Slack message starting with “@channel urgent” every time someone deploys.

The Intern, the BFF, and the SPA

Let me give you the full picture of that war room incident, because the details matter more than the headline.

We had some interns working on one of our BFFs. The client code and the BFF lived in the same repo — a mono-repo setup that’s common at Agoda and plenty of other places. The intern looked at this and made a perfectly logical assumption: since the client and the server are in the same repo, they deploy together. So a breaking change to the API would be fine, right? The client code would just update at the same time.

They were young. It’s a reasonable mental model if you’ve never worked with SPAs before. A single-page application doesn’t redeploy when your server does. The JavaScript your users downloaded an hour ago — or yesterday, or last week if they haven’t refreshed — is still calling the old contract. Your server moved; your client didn’t. It’s the same repo, but it’s not the same deployment.

What makes this story worth telling isn’t the intern’s mistake — it’s that it was missed in review too. Experienced engineers looked at the diff, saw the improvement, and approved it. The breaking change was done the right way in every other sense: it was behind an experiment, which meant we didn’t need to roll back the deployment. We just killed the experiment. But the spike in 4xx responses was real, the war room was real, and the lesson was real.

The gap wasn’t knowledge. It was that we hadn’t built the instinct — the reflex — to ask: “Can an existing client keep working without modifying their code?” That single question, asked every time, is worth more than any API design document you’ll ever write.

Your API Deserves the Same Thought as Your Architecture

Here’s the uncomfortable truth about API design: you get one chance to get the contract right, and most teams spend more time debating variable names in a pull request than they do thinking about the shape of a response that three teams will depend on for the next three years.

We plan our database schemas. We whiteboard our system architecture. We argue about domain boundaries for weeks. But the API — the actual surface that consumers will build against, the thing that has to remain stable while everything behind it evolves — that gets designed in the time between standup and lunch by an engineer thinking about the feature they need to ship, not the contract they’re about to publish.

The physicist Richard Feynman once said, “The first principle is that you must not fool yourself — and you are the easiest person to fool.” We fool ourselves into thinking API design is just implementation that happens to be public. It isn’t. Implementation you can refactor on Tuesday. An API contract, once someone builds against it, is a conversation — and breaking changes are conversations you have at 2 AM on Slack with a very unhappy team.

Consistency is what makes an API learnable. When every endpoint follows the same conventions, engineers stop reading documentation and start predicting behaviour. They know where the resource will be, how errors will look, what happens when they send a bad request. That predictability isn’t a nice-to-have — it’s the difference between an API that people integrate with confidently and one that requires a Slack message to the owning team for every new endpoint. Getting this right upfront is dramatically cheaper than fixing it after consumers have already built against your inconsistencies.

So before we get into the specifics of what a good API contract looks like, let’s agree on one thing: this deserves deliberate thought. Not a design committee. Not a twelve-page RFC. Just the same care you’d give to any other architectural decision that will outlive your current sprint.

Resource Naming: Let HTTP Do the Verbs

The simplest API design rule has the biggest impact, and yet we keep getting it wrong. Let HTTP methods handle the verbs. Let your URLs describe the thing.

Don’t build GET /getBooking?id=123. Build GET /bookings/123. Don't build POST /createRatePlan. Build POST /rate-plans. The URL describes the resource; the method is the action. POST already means "create." PUT already means "update." You don't need to say it twice.

Use plural nouns for collections (/bookings, /rate-plans). Use hierarchy to show ownership (/properties/{id}/rate-plans). Pick a casing convention and stick with it — we're not going to relitigate camelCase versus kebab-case here, but inconsistency across endpoints is worse than either choice. Use query parameters for filtering and sorting, not for identity (/bookings?status=confirmed&sort=checkIn).

If you want to see what “doing it right” looks like at scale, look at Stripe’s API. It’s almost universally cited as the gold standard in API design, and for good reason. Every resource follows the same conventions. Every endpoint behaves predictably. New engineers can guess the URL for an endpoint they’ve never used because the patterns are consistent. That didn’t happen by accident — it’s the result of a deliberate design review process that treats the API as a product, not a side effect of implementation. Brandur Leach wrote about their approach in detail back in 2017, and it’s still the benchmark a decade later.

One thing we’ve been historically good at here at Agoda — almost accidentally, if I’m being honest — is using business language in our APIs. If the team says “rate plan,” the URL says /rate-plans. If the domain calls it a "booking," nobody's out there naming it /reservations because that's what they called it at their last company. This is the DDD concept of ubiquitous language applied to API design, and when it works, it eliminates an entire category of "wait, what does this endpoint actually do?" conversations.

Versioning: You Probably Don’t Need It

Most versioning debates I’ve sat through focus on the mechanism — URL path versus header — while ignoring the more important question: do you actually need a version bump at all?

Adding an optional request field with a sensible default? No version bump needed. Adding a new field to a response, assuming your consumers ignore unknown fields? No version bump. Adding an entirely new endpoint? Still no. But removing a field, renaming a field, changing a type, tightening validation that previously accepted something? Those are breaking changes. Those need a version or a deprecation plan or both.

The pattern is straightforward: additive changes are usually safe; subtractive or mutating changes are breaking. If you design for additive evolution from the start, you rarely need version bumps at all.

I always tell my engineers the same thing: “Don’t do breaking changes.” And they look at me like I’ve just asked them to write code without an IDE. So let me break this down. I contribute to NUnit when I’m not busy — which admittedly means it’s been a few years since my last PR — but NUnit went roughly nine years between major versions (version 3.0 shipped in 2015, version 4.0 in 2024). Now, NUnit is a library, not a BFF. But the principle is the same, just amplified: if an open-source project with thousands of consumers across the .NET ecosystem can go nine years without a breaking change, I’m fairly confident you can manage it with an internal API that has three consumers.

When you genuinely do need a version, start at /v1/. Evolve in place with additive changes. Reserve /v2/ for real, unavoidable breaking changes. Most services never need v2 if they were designed for evolution. Our APIs are about 90% internal — team-to-team — which gives us more flexibility than external-facing APIs. But "internal" doesn't mean "casual." Internal consumers deserve the same stability as external ones. They just can't leave you a one-star review.

The 200 OK That Lies to Your Face

I need to talk about an anti-pattern that will resonate with every engineer who’s ever debugged a “successful” API call that actually failed.

By default, a Sangria GraphQL server will return HTTP 200 even when the query has execution errors, as long as the request itself was valid at the transport level. This follows the GraphQL spec convention, which treats GraphQL execution errors differently from transport errors. And on paper, that distinction makes sense. In practice, it means your monitoring thinks everything is fine, your load balancer thinks the service is healthy, your retry logic never triggers, and your engineers are left staring at a 200 OK response that contains, buried in the JSON, a field quietly announcing that everything went wrong.

The coach Vince Lombardi said, “Practice does not make perfect. Only perfect practice makes perfect.” We’ve been perfectly practising the wrong thing with our error contracts — returning success codes for failures so consistently that we’ve trained our entire monitoring infrastructure to miss real problems.

A good error response has three parts. The HTTP status code gives you the category — 4xx for client errors, 5xx for server errors. A stable, machine-readable error code gives your consumers something to program against — RATE_PLAN_NOT_AVAILABLE is infinitely more useful than parsing a free-form string that someone will reword next sprint and break your handling. And a human-readable message gives engineers and support something to read in logs.

HTTP/1.1 422 Unprocessable Entity
{
  "error": {
    "code": "RATE_PLAN_NOT_AVAILABLE",
    "message": "The selected rate plan is no longer available for the requested dates.",
    "details": {
      "rate_plan_id": "rp_2847",
      "requested_check_in": "2026-03-15"
    }
  }
}

Compare that to what we’ve all seen in the wild:

HTTP/1.1 200 OK
{
  "success": false,
  "message": "Something went wrong, please try again later."
}

The first one tells you what happened, gives your code something stable to switch on, and provides enough context to debug without digging through logs. The second one lies to your monitoring with a 200 status, gives your code nothing to work with, and tells the engineer exactly nothing about what went wrong. One error format per API. Documented. Treated as part of the contract. An undocumented error code is like an undocumented API endpoint — it exists, people depend on it, and it’ll break something when you change it.

Backward Compatibility as a Discipline

Everything above — naming, versioning, error contracts — serves one goal: your consumer can keep working without a code change when you deploy.

That’s it. That’s the test. Before every change, ask: “Can an existing client keep working without modifying their code?” If the answer is no, it’s a breaking change. Full stop. It needs a version bump, a deprecation plan, or a conversation.

The deprecation workflow we follow isn’t glamorous, but it works. Announce the deprecation with a timeline — “v1 sunset in 90 days.” Add a Deprecation header so consumers can detect it programmatically. Provide a migration guide that's actually useful, not a changelog formatted as a guide. Monitor usage — reach out to teams still on the old version. Then, and only then, sunset it.

The architect Christopher Alexander wrote, “When you build a thing you cannot merely build that thing in isolation, but must repair the world around it, and within it, so that the larger world at that one place becomes more coherent.” We’ve been doing the opposite with our APIs — building things that make the world around them less coherent, one breaking change at a time. Every field rename, every tightened validation, every removed endpoint is a small fracture in the trust that lets teams deploy independently.

If you’ve read my post on backwards compatibility testing patterns, you’ll know we’ve built tooling at Agoda to catch these issues before they escape — compiling against both the new interfaces and the last published version to make the build system enforce what code review sometimes misses. That’s the safety net. But the real discipline happens earlier, in design, when you decide to add an optional field with a default instead of renaming the existing one.

The Five-Minute Code Review Checklist

If I’ve done my job with this post, you should be able to run through these checks on any API change in under ten minutes:

Does the URL describe a resource with a plural noun, or is it a verb disguised as an endpoint? Is this an additive change, or does it remove, rename, or mutate something existing? If it’s breaking, is there a versioning and deprecation plan? Does the error response use proper HTTP status codes, or are we hiding failures behind 200 OK? Can an existing client — including one running JavaScript downloaded three days ago — keep working without a code change?

That last question is the only one that really matters. The rest are just ways of getting there.

The Bottom Line

The teams that build APIs worth depending on aren’t doing anything magical. They’re asking one question, consistently, at every boundary: “Can my consumers keep working when I deploy?” They’re treating their API as a product with its own design standards, not as a side effect of their implementation. They’re designing for additive evolution so they rarely need breaking changes. And when they do need to break something, they’re giving consumers a migration path that respects their time.

As I’ve written before, the discipline of library APIs and service APIs is the same discipline — keep your promises, evolve without breaking, and treat every public surface as a contract that someone’s production system depends on. Your code entropy grows fastest at the boundaries between systems, and APIs are the biggest boundaries you have.

Now, if you’ll excuse me, I need to go review a PR where someone’s proposing to “just rename” a response field on one of our BFFs. The field name is genuinely terrible — I think it was named during a Friday afternoon session that may or may not have involved beer — but there are browser tabs out there right now parsing that field, and every one of them is a promise I intend to keep.