Resilience Isn’t a Feature — It’s the Absence of Naivety
Or: Why Your System Is One Slow Query Away from a Very Bad Night
The Grafana dashboard had been green all day. At 2:14 AM on a Thursday, the on-call engineer’s phone lit up the bedside table — not one alert, but eleven, arriving in a cascade that sounded like a slot machine paying out in the worst way possible. He sat up, squinted at the screen, and watched three services go red in the time it took to unlock his laptop. The vinyl floor was cold under his feet as he walked to the kitchen for water he wouldn’t drink. By the time he’d opened the VPN, the fourth service was down.
That’s not a hypothetical. That’s what happens when distributed systems discover, at the worst possible moment, that nobody taught them how to fail.
Every microservices tutorial on the internet will teach you how to make services talk to each other. Almost none of them cover what happens when they stop. This is like teaching someone to drive by explaining the accelerator and the steering wheel but skipping the brakes, the seatbelt, and what to do when the engine catches fire on the expressway.
As the management theorist Peter Drucker put it, “The greatest danger in times of turbulence is not the turbulence — it is to act with yesterday’s logic.” In distributed systems, yesterday’s logic is the assumption that your dependencies will be there when you need them. Today’s reality is that at any given moment in a system with fifty services, something is slow, something is timing out, something is returning garbage data, and something is completely down. The question isn’t whether your system handles failure. It’s whether it handles failure gracefully — or whether it handles it by propagating the problem to everything else until the whole thing collapses like dominoes on a wobbly table.
The Anatomy of a Cascade
Let’s walk through what actually happens, because understanding the mechanics is the first step toward building the instinct to prevent it.
Service A calls Service B. Service B calls Service C. Service C calls a database. The database is slow — maybe someone ran an analytics query, maybe a table lock went sideways, maybe it’s Tuesday and Tuesdays are just like that sometimes. The query that normally returns in 40 milliseconds is now taking 90 seconds.
Service C’s threads are all waiting. Its thread pool exhausts. Service B’s threads are all waiting on Service C. Its pool exhausts. Service A’s threads are all waiting on Service B. Its pool exhausts. The load balancer starts returning 503s. Users see error pages. Your monitoring lights up like a Christmas tree. And somewhere in the middle of it all, a single slow database query is sitting there, blissfully unaware that it just brought down three services and ruined several thousand people’s evening.
“For want of a nail the shoe was lost. For want of a shoe the horse was lost. For want of a horse the rider was lost.”* — Proverb, dating to the 13th century*
That isn’t a metaphor. That is literally how cascading failures work. A missing nail — a single slow dependency — collapses the entire chain because nothing in it knows how to fail gracefully.
The Three Things Every External Call Needs
We’re not talking about advanced patterns here. We’re talking about the basics — the seatbelts and brakes that should be on every service before it goes anywhere near production.
Timeouts: Stop Waiting for Godot
Without a timeout, a slow dependency turns your fast service into a slow service. Your threads block. Your pool exhausts. Your callers block. The cascade begins.
Every external call needs a timeout. HTTP calls. Database queries. gRPC calls. Cache lookups. Message queue operations. No exceptions.
Here’s the thing that still surprises me after years of doing this: the default timeouts on most HTTP and SQL clients are insane. We’re talking minutes. Minutes. In what universe does waiting two minutes for a response represent a functioning system? At Agoda, the first thing we change out of the box is setting timeouts to milliseconds. If you haven’t responded in less than a second, something is wrong, and I’d rather fail fast, forget about it, and retry the next node than sit there holding the line while the whole system backs up behind me.
A timeout is actually an optimistic pattern, if you think about it. It says “I believe this call will succeed — but if it doesn’t respond quickly, something is wrong and I’d rather know now.” Without a timeout, you’re saying “I’ll wait as long as it takes,” which in a distributed system can mean forever. And forever is a long time to block a thread.
Retries: The Second Chance (Done Right)
The naive approach — retry immediately on failure — is the pattern that turns a struggling service into a dead service. If a service is slow because it’s overloaded, hammering it with retries is like seeing someone drowning and throwing them more water.
We learned this the hard way at Agoda. We had an incident where a service five levels deep in the call stack started timing out. Every layer above it had retries configured. The web browser had retries. The BFF had retries. The mid-tier services had retries. When that bottom service got slow, it didn’t get fewer requests to help it recover — it got exponentially more requests as every layer above it dutifully retried. The system that was under pressure got crushed under the weight of everyone trying to help.
After that incident, we started adding and propagating headers with retry counts so that downstream systems could see how many times a request had already been retried and react accordingly. If you’re the fifth retry of a request that’s already been retried by three layers above you, maybe the kindest thing you can do is say no.
The evolution of retry strategies tells you everything about how we’ve learned to respect distributed systems:
No retry means transient errors cause unnecessary failures. Immediate retry means you hammer the dying service. Exponential backoff gives the service time to recover. Backoff with jitter prevents the thundering herd — that terrifying phenomenon where a thousand clients all retry at exactly the same backoff interval, creating a synchronised retry storm that’s worse than the original load.
Circuit Breakers: Knowing When to Stop Trying
Your house has a circuit breaker. When the electrical current exceeds safe levels, the breaker trips and cuts the circuit. It doesn’t keep pushing current through and hope the wiring doesn’t catch fire. It stops and waits.
Software circuit breakers work the same way. When a dependency’s failure rate exceeds a threshold, the circuit opens and requests fail immediately — you don’t even try. After a timeout period, you let a few requests through to see if the service has recovered. If it has, the circuit closes and normal operation resumes. If it hasn’t, the circuit stays open.
Here’s the emotional insight that most engineers miss: failing fast is an act of kindness to the rest of the system. By not calling a broken service, you’re protecting your own resources, protecting the broken service from additional load it can’t handle, and giving your users a fast error instead of a slow timeout. A quick “no” is almost always better than a long silence.
Bulkheads: Containing the Blast Radius
The name comes from the watertight compartments in ships. The Titanic sank because water flowed between compartments. In software, without bulkheads, one failing dependency drowns your entire thread pool.
Without isolation, all your threads live in one pool. A slow call to Service B consumes threads. Perfectly healthy calls to Service A can’t get threads because they’re all stuck waiting on Service B. With bulkheads, each dependency gets its own pool. Service B’s pool exhausts, but Service A’s pool is untouched and keeps working.
Netflix built their entire Hystrix library because of exactly this problem. When one of their hundreds of services slowed down, it would consume the thread pools of its callers, which consumed the thread pools of their callers, until the entire system collapsed. Hystrix introduced bulkheads, circuit breakers, and fallbacks as first-class architectural concerns — not afterthoughts bolted on after the third outage.
At Agoda, we took a different path to the same destination. Over a decade ago, we had a massive outage caused by a load balancer failure. That experience burned us deeply enough that we removed traditional load balancers entirely and moved to client-side round robin with service discovery. Our teams wrote custom HTTP and SQL clients that handle resilience at the client level — service discovery, round robin, retries, timeouts, all baked in. Not only did this eliminate that single point of failure, it saved us significant infrastructure cost. These libraries are community-maintained across the engineering organisation, and we’ve even open-sourced some of them.
Recently, we’ve been moving toward Envoy for some of this, but the principle remains: resilience lives in the client, not in a shared piece of infrastructure that becomes everyone’s problem when it goes down.
Fallbacks: The Business Decision Nobody Made
This is where the conversation shifts from engineering patterns to leadership responsibility, and where I need every engineering manager and product owner reading this to pay attention.
When a service is down, you have options. You can return cached data — show yesterday’s prices. You can return a default — “Price unavailable, call for a quote.” You can degrade gracefully — hide the reviews section entirely. Or you can fail fast and show an error page.
Here’s the problem: the choice between these options is a business decision, not an engineering decision. Should you show stale prices or no prices? If you’re running a hotel booking site, showing yesterday’s price might lead to a booking at the wrong rate — that’s a financial cost. But showing no price means the user leaves — that’s a revenue cost. Which is worse?
Your engineer should not be making this call alone at 3 AM with one eye open and cold coffee going stale on the desk. Your product owner should have made it at 3 PM last Wednesday, and it should be documented and implemented as a fallback strategy before the failure happens.
“By failing to prepare, you are preparing to fail.”* — Benjamin Franklin*
Franklin wasn’t thinking about microservices, but he nailed it anyway. Resilience isn’t something you bolt on during the incident. It’s a conversation you have before the incident, between product and engineering, about what “good enough” looks like when “perfect” isn’t available.
Something we do at Agoda that I think is genuinely unique is feature shedding. When traffic spikes beyond what the system can comfortably handle, we start deliberately dropping compute-expensive features. Not randomly — strategically. We identify which features are expensive to render and which are critical to the core user journey, and when the system is under pressure, we shed the expensive non-critical ones. The user still gets a functioning booking experience. They just might not get personalised recommendations or that fancy interactive map while the system is under stress. It’s the difference between dimming the lights to keep the power on and letting the whole grid go dark.
The Real Cost of Not Being Resilient
Let’s do the maths, because this is where abstract patterns become concrete urgency.
Service A calls Service B calls Service C. Without timeouts, if C takes 60 seconds, B waits 60 seconds, A waits 60 seconds plus B’s processing time. If A has 100 threads and C goes slow, all 100 threads block within minutes. Service A is now effectively down — not because anything is wrong with A’s code, but because it trusted C unconditionally.
Amazon’s 2017 S3 outage is the canonical example. A simple operational error in S3 caused failures across dozens of AWS services because so many services depended on S3 without adequate resilience patterns. The blast radius extended far beyond what anyone anticipated because the dependency was so deeply embedded in everything. One service. Hundreds of dependents. Hours of cascading failure.
Michael Nygard’s book Release It! frames this perfectly: stability patterns exist because instability patterns are the default. In a distributed system, the absence of resilience patterns isn’t a neutral state. It’s an active risk. You’re not being conservative by skipping the circuit breaker. You’re being reckless.
Be Your Own Chaos
You might be wondering whether we do chaos engineering at Agoda — intentionally breaking things to test resilience. The honest answer is: not formally. We are our own chaos.
With a culture of ruthless experimentation and thousands of A/B tests running simultaneously, we create enough organic turbulence to keep our systems honest. We do scheduled traffic shifts to test the capacity of our fallback paths, and the feature shedding I mentioned earlier gets exercised more often than you’d think. But we haven’t needed to inject artificial chaos because the combination of scale, experimentation, and continuous deployment provides a steady stream of real-world failure scenarios.
That said, this isn’t a recommendation to skip chaos engineering. It’s an observation that if you’re running at sufficient scale with sufficient experimentation velocity, your system is already being stress-tested constantly. The question is whether you’re paying attention to what those tests are telling you.
Resilience as Culture, Not Configuration
Here’s where we land, and it’s the point I want to leave you with: resilience isn’t a sprint feature. You don’t ship it in a two-week cycle and move on. It’s a design posture — a default way of thinking about every external call, every dependency, every integration point.
The engineers who build resilient systems aren’t the ones who memorised the patterns. They’re the ones who assume every external call will fail and design accordingly. That assumption changes everything about how you write code, how you review code, and how you think about architecture.
“Where’s the timeout?” should be as automatic in code review as “where’s the null check?” “What happens when this dependency is down?” should be the first question in every architecture discussion, not the last. “What’s the fallback?” should be in every feature specification. “Why didn’t we fail gracefully?” should be a standard line item in every post-incident review.
“Everyone has a plan until they get punched in the mouth.”* — Mike Tyson*
Every service has a happy path. Resilience is what happens when the happy path gets punched in the mouth. And the time to plan for that punch is now — not when you’re lying on the canvas at 2 AM, watching your Grafana dashboard turn red one panel at a time.
Now, if you’ll excuse me, I need to go check the timeout configuration on that new service we deployed yesterday. I’m sure it’s fine. It’s always fine. Until it isn’t.