Semantic Monitoring: The Question You’re Not Asking About Your Production Systems
Or: Why Your Dashboards Are Green and Your Business Is Bleeding
There’s a question that separates engineers who truly own their systems from those who merely operate them: “What does this application do?”
Not what technologies it uses. Not what endpoints it exposes. Not what its CPU utilisation looks like at 3 PM on a Tuesday. But what does it actually do — what value does it deliver to the humans who depend on it?
If you can’t answer that question in a single sentence, you’re already in trouble. And if your monitoring isn’t built around that answer, you’re flying blind with your eyes wide open, staring at dashboards full of green lights while your business quietly bleeds out.
As Warren Buffett’s longtime partner Charlie Munger once observed about academics: “They studied what was measurable, rather than what was meaningful.” We’ve built entire observability stacks on that exact mistake.
The “You Build It, You Run It” Starting Point
Before we dive into specifics, let’s establish the philosophy that makes any of this work. Here at Agoda, we follow the principle of “You Build It, You Run It.” This isn’t just a catchy DevOps slogan — it’s a fundamental shift in accountability. When you’re the one getting woken up at 3 AM, you develop a visceral understanding of what actually matters in production.
And here’s what that experience teaches you: most of what we monitor is noise. Useful noise, sometimes. Diagnostic noise, occasionally. But noise nonetheless. The signal — the thing that tells you whether your system is actually doing its job — is almost always buried somewhere else entirely.
What Can We Actually Monitor?
Let’s be honest about our options. When it comes to observing our systems, we typically focus on three categories:
Logs give us the narrative — application events, web server activity, system-level happenings. They’re the “what happened” of your system.
Metrics give us the numbers — performance indicators, resource utilisation, request counts, response times. They’re the “how much” of your system.
Traces give us the journey — the path a request takes through your distributed system, where it spent its time, and where things went wrong. They’re the “how did we get here” of your system.
All useful. All important. But here’s the uncomfortable truth: you can have perfect scores across all three categories and still have a product that’s failing its users. Your servers can be humming along beautifully while your customers are abandoning their shopping carts in frustration.
This is where we need to talk about semantic monitoring.
The Top-Down View: What Does Your Application Actually Do?
Sam Newman, in his book Building Microservices, introduces the concept of semantic monitoring, which is closely related to (and sometimes overlaps with) what is often called synthetic monitoring. The idea is deceptively simple: instead of just monitoring whether your system is running, monitor whether it’s doing what it’s supposed to do.
The “semantic” in semantic monitoring refers to meaning. You’re not just checking that an HTTP request returns a 200 OK — you’re checking that the meaning of your system’s purpose is being fulfilled. A 200 OK response, after all, is not always OK. The request might have succeeded technically while failing completely at delivering value.
Let me give you a concrete example. If you’re running an e-commerce website, what does your application do? Strip away all the microservices, the deployment pipelines, the Kubernetes clusters. At its core, your system exists to facilitate orders. That’s it. That’s the ultimate goal — guiding users through the funnel to complete a purchase.
So how do we use this insight for monitoring? We ask two brutally simple questions:
Is the application successfully processing orders?How often should it be processing orders, and is it meeting that frequency right now?Everything else is context. Important context, diagnostic context — but context nonetheless.
Dynamic Baselines: Making Semantic Monitoring Practical
Here’s how we’ve implemented this at Agoda. Rather than setting static thresholds (which are either too sensitive during quiet periods or too lenient during busy ones), we look at 10-minute windows throughout the day. For each window, we calculate the average number of orders for that same time slot, on the same day of the week, over the past four weeks.
If current performance falls significantly below this baseline, it triggers an alert.
This approach is more nuanced than it might first appear. It naturally adapts to the ebb and flow of traffic throughout the day and week. You’re not getting false alarms at 4 AM when volume is legitimately low, and you’re not missing genuine problems at 4 PM when you’d normally be processing thousands of transactions.
The dangerous alternative is the static threshold trap. To avoid constant alerts during off-hours, teams often set thresholds embarrassingly low. But that same low threshold means you might miss a 50% drop during peak hours — the exact moment when it matters most. Our dynamic approach solves this by adjusting expectations based on historical patterns.
You can tune the window size to your context. High-volume systems might use 10-minute windows. Lower-volume systems might need hourly or even daily windows. Yes, this means slower response times for lower-volume systems — but you’ll still catch significant issues, and you won’t be drowning in false positives.
And I manage some systems at Agoda that use daily windows, it takes us a day to know we have a problem, but when we do wee know. Sometimes when you are small you don’t have the volumes, but that’s ok, it can still work.
Another Example: The Help Desk Chat System
Let’s apply our core question to a different type of application: a help desk chat system. What does this system do?
At its most basic level, it allows communication between support staff and customers. But let’s break that down further:
It enables sending messagesIt displays these messages to the participantsIt presents a list of ongoing conversationsNow, you might be tempted to say that sending messages is the primary function. And you’d be partly right. But remember, we’re thinking about what the system does, not just how it does it. The system exists to help customers solve their problems efficiently.
With this in mind, what should we monitor?
Tracking successful message sends is important, but it might not tell the whole story — especially if message volume is low. We should also consider:
Successful page loads for the conversation list — Can users see their ongoing chats with an increasing number of new messages? think seen tracking for example.Successful loads of the message window — Can users access the core chat interface?Successful resolution rate — Are chats actually leading to solved problems?By expanding beyond “are messages sending?” we get a comprehensive view of whether the system is truly fulfilling its purpose: helping customers solve their problems efficiently.
This is why the question “What does this system DO?” is so powerful. It guides us towards monitoring metrics that reflect the health and effectiveness of our product, not just its technical performance.
A 200 OK Is Not Always OK
Here’s a truth that took me years to fully internalise: successful technical operations are necessary but insufficient for product health.
Your database query might execute perfectly and return exactly what it was asked for — while returning data that’s three hours stale because a synchronisation job silently failed. Your API might respond with a textbook 200 OK while serving up a response that contains subtly corrupted data. Your message queue might be humming along with zero errors while messages pile up faster than they’re being processed. Or have you ever seen an empty catch statement? they exist too.
The technical layer doesn’t understand business context. It can’t. That’s not its job. Its job is to tell you whether operations succeeded or failed at the infrastructure level. Your job is to layer meaning on top of that.
The Bottom-Up View: How Does Your Application Work?
None of this means traditional monitoring is useless — far from it. The bottom-up approach looks at the internal workings of your application: HTTP requests and their response times and codes, database calls and their success rates, message queue depths and processing rates.
Modern systems often collect these metrics through contactless telemetry or auto-instrumentation, reducing the need for custom instrumentation. This is genuine progress. The plumbing of observability has never been easier to install.
But here’s the key insight: bottom-up monitoring tells you how your system is operating. Top-down semantic monitoring tells you whether your system is fulfilling its purpose. You need both, but you need to know which one answers which question.
Prioritising Alerts: The 3 AM Question
This brings us to one of the most politically charged questions in any engineering organisation: when should someone get woken up?
Ask yourself this: Should the Network Operations Centre call you at 3 AM if a server hits 100% CPU usage?
The answer is no — not if there’s no business impact. If your core business functions (like processing orders) are unaffected, it’s better to wait until the next day to address the issue. CPU usage is a diagnostic metric, not a semantic one.
This is where semantic monitoring becomes not just a technical approach, but a sanity-saving one. When your alerts are tied to business outcomes, you can sleep through the infrastructure hiccups that aren’t actually hurting anyone. When they’re tied to resource utilisation, every spike becomes a potential crisis.
Using Loss as a Currency for Prioritisation
Once you’ve established a health metric for your system and can compare current performance against your baseline, you unlock something powerful: the ability to quantify “loss” during a production incident.
Imagine your e-commerce platform typically processes 1,000 orders per hour during a specific time window, based on your four-week average. During an incident, this drops to 600 orders. You can now quantify your loss: 400 orders per hour. If you know your average order value, you can translate this into actual revenue impact.
This quantification becomes your currency for making critical decisions. You can compare the impact of multiple ongoing issues. You can justify allocating more resources to high-impact problems. You can make data-driven decisions about when it’s actually worth waking up engineers in the middle of the night.
Reid Hoffman, co-founder of LinkedIn, captured this perfectly: “You won’t always know which fire to stamp out first. And if you try to put out every fire at once, you’ll only burn yourself out. That’s why entrepreneurs have to learn to know which fires they can let burn — and sometimes even very large fires.”
This wisdom applies directly to our world of production incidents. Sometimes, you have to ask not which fire you should put out, but which fires you can afford to let burn. Your loss metric gives you a clear way to make these tough decisions.
Beyond Incident Response
The concept of loss as currency extends beyond immediate firefighting. You can use it to prioritise your backlog, make architectural decisions, and guide your product roadmap. When you propose investments in system improvements or additional resources, you can back those proposals with clear figures showing the potential loss you’re trying to mitigate.
Yes, there’s some crystal-ball gazing involved in estimating how likely incidents are to recur. But a rough estimate grounded in business impact beats a precise measurement of something that doesn’t matter.
By always thinking in terms of potential loss (or gain), you ensure that your team’s efforts align with what truly matters for your business and your users. You create a direct link between your technical decisions and your business outcomes.
The Bottom Line
As we’ve explored throughout this post, measuring product health goes far beyond monitoring code quality or individual system metrics. It requires a holistic approach that starts with a fundamental question: “What does our system DO?”
This simple yet powerful query guides us toward understanding the true purpose of our products and how they deliver value to users. By focusing on core business metrics that reflect this purpose, we can create dynamic monitoring systems that adapt to the natural rhythms of our product usage.
The goal isn’t just to have systems that run smoothly from a technical perspective. It’s to have products that consistently deliver value to our users and meet our business objectives. Semantic monitoring makes this possible by shifting our focus from “is it running?” to “is it working?”
Remember: you can’t measure what you’re not monitoring, but you also can’t improve what you’re measuring wrong. The metrics that matter are the ones that tell you whether your system is fulfilling its purpose — not just whether its components are functioning.
As you apply these principles to your own systems, always start with that core question: “What does this system DO?” Let the answer guide your metrics, your monitoring, and your decision-making. In doing so, you’ll not only improve your product’s health but also ensure that your engineering efforts are always aligned with what truly matters for your business and your users.
Now, if you’ll excuse me, I need to go check why our order volume is down 15% in the last hour. The dashboards all look green, but the business metric tells a different story.