Experimentation and Risk Taking: Building a Culture That Bets on Better
I’ll never forget the day our CEO made a simple bet. We were in our monthly product meeting, and one of our product owners was presenting what everyone in the room knew was a genuinely terrible idea. You could feel the collective cringe. But instead of shutting it down, our CEO leaned back and said, “I’ll bet you a cup of coffee that doesn’t work.”

The message was crystal clear: without data to back it up, that’s all any of our opinions were worth — a cup of coffee. He wasn’t being dismissive; he was establishing the culture we needed.
You see, we engineers have this beautiful curse — we fall in love with our solutions. We craft elegant algorithms, we optimize performance to the millisecond, and we convince ourselves that because the code is beautiful, the impact will be too. But as Maya Angelou once said, “When people show you who they are, believe them.” Your users will show you what they actually want through their clicking, scrolling, and converting — believe that data.
The Foundation: Building Your Data Platform
Before you can start making informed bets, you need the infrastructure to measure them. The good news? We live in an era where this has never been easier. Firebase offers built-in A/B testing capabilities that you can implement in hours, not months. Google BigQuery can handle your experimentation data at scale, and honestly, for most scenarios, you could roll your own A/B testing framework in a weekend.
The key is logging everything from day one. When you implement two different code paths — one for your control group, one for your experiment — you’re essentially creating a time machine that lets you see what would have happened if you’d made a different choice.
When NOT to Use A/B Testing
But here’s where it gets interesting. A/B testing isn’t always the answer, and knowing when to skip it is just as important as knowing when to use it.
When you’re small and the impact is obvious: If you’re a startup with 100 users and you’re testing whether to fix a broken login button, you don’t need statistical significance. You need that button to work. Period.
When the change is night and day: Sometimes you know a change will be transformative. In the very early days of our extranet portal, we introduced a simple CTA (call to action) that let users bulk add promotions to multiple products, instead of redirecting them to our complex promotions page for each individual product. We didn’t A/B test it because we knew forcing users through a tedious multi-step process was killing adoption. We went from 2% daily promotion activation to 30% overnight. Sometimes the improvement is so obvious that you just need to ship it.
Library upgrades and dependency changes: This is where it gets tricky. Should you A/B test upgrading from one version of React to another? What about switching from NuGet packages to bundled dependencies? At Agoda, we’ve experimented with everything from bundling strategies to completely different tech stacks like Zypher for specific services. The key is understanding what you’re actually measuring.
The Small Changes Paradox
Here’s something that might surprise you: small changes should be the easiest to A/B test, yet they’re often the ones we skip. Adding a single line of CSS? Changing button text? These feel insignificant, but they’re perfect candidates for experimentation because the implementation cost is minimal and the learning is pure.
“In God we trust. All others must bring data.” — W. Edwards Deming
Even bug fixes can be worth testing. I know it sounds counterintuitive — surely fixing bugs is always good? — but by A/B testing bug fixes, you can demonstrate their business value to stakeholders. When we fixed a critical bug in our booking flow and showed that it directly led to a 12% increase in completed bookings, suddenly quality became a C-suite conversation.
Testing Culture vs. A/B Testing Culture
Notice a pattern in those examples above? Most of the times we skip A/B testing is when we’re small. And that’s exactly where you should start — with a testing culture, not an A/B testing culture.
When you’re small, focus on simply measuring impact. Did that new feature increase signups? Did fixing that bug reduce support tickets? Did simplifying that workflow improve conversion? You don’t need statistical significance when you have 100 users — you need to build the habit of measuring everything.
This is crucial: if you get into this measurement mindset early, adding proper A/B testing infrastructure later becomes a business necessity, not something you have to convince people to invest in. Your stakeholders will already be addicted to data-driven decisions.
Connecting Experiments to North Star Metrics
Here’s where many teams get seduced by vanity metrics. Your login redesign increases signups by 500%? That sounds incredible, but what if those new users never actually purchase anything? What if the simplified signup process is attracting low-intent users who dilute your conversion funnel?
This is why connecting your measurements to north star metrics — like actual ecommerce sales, revenue per user, or customer lifetime value — becomes critical as you mature. You might celebrate that signup boost while your sales team watches conversion rates plummet.
The progression every engineering organization needs to make is moving from measuring everything to measuring what matters. Don’t only micro-measure every button click and page view — you still need those metrics, but it’s the big ones that move the dial for your business. Instead, draw clear lines between your experiments and business outcomes. That login change should ultimately connect to revenue, not just user registrations.
This business value connection transforms how stakeholders view engineering experiments. When you can show that your A/B test directly contributed to a 15% increase in quarterly revenue, suddenly everyone wants to know what you’re testing next.
What Experimentation Looks Like in Practice
Let me give you some real examples of how we approach this:
Technical Implementation Testing: We had two engineers arguing about the best way to implement local caching. Instead of endless architecture debates, we implemented it both ways and ran them as experiments in production, measuring performance differences. We also tracked conversion rates throughout the funnel to ensure our technical experiment didn’t negatively impact the business — if anything, we usually see slightly positive results from performance gains. Is this failing? In a way, yes — one approach had to lose. But we learned which solution actually worked better under real conditions, not theoretical ones. We’ll cover this approach to failure in more detail in our next post.
Product Feature Testing: Adding review capabilities to our product pages illustrates a key experimentation challenge. If you just launch an A/B test showing empty review sections, you’ll probably hurt conversion — seeing no reviews makes users nervous, and empty review sections across every product would look broken. In these scenarios, we’ve had to get creative: sourcing initial reviews from third parties, or running the collection experiment separately from the display experiment. Sometimes it isn’t as simple as “turn everything on and measure.” You need to plan your experimentation so what you’re testing on users is the experience you want, not just the feature you’re building.
UX/Design Changes: Every interface change gets the A/B treatment. Button colors, form layouts, even the positioning of error messages. You’d be amazed how often your intuition is wrong. This is where it will really shine for you.
When Experiments Fail (And Why That’s Beautiful)
“Success is not final, failure is not fatal: it is the courage to continue that counts.” — Winston Churchill
The most valuable experiments are often the ones that fail. We ran experiments to make our mobile layout more consistent with our desktop layout — it would make development easier, and we thought users would appreciate the consistency. Less confusion, right? Wrong. Our users who stick with mobile expect mobile patterns, and desktop users expect desktop patterns. The changes just confused everyone. In one instance, we moved the product image further down the mobile page to match our desktop hierarchy where navigation elements come first. The conversion loss was brutal. Our UX consultant had to explain the obvious: users love images, especially on mobile where real estate is precious. Every pixel needs to drive conversion, not match some arbitrary cross-platform consistency that only engineers care about.
This is where betting culture becomes crucial. When your CEO says “I’ll bet you a cup of coffee that doesn’t work,” they’re encouraging you to prove them wrong with data, but at the same time, if you fail, there’s a learning there we can take. They’re creating psychological safety around failure.
Building a Betting Culture
The goal isn’t to be right all the time. The goal is to be learning all the time. When someone on your team says “Let’s try this,” your response shouldn’t be “Are you sure?” It should be “How will we measure it?”
At Agoda, we’ve fostered this by celebrating failed experiments just as much as successful ones. Because every failed experiment is just a successful elimination of a bad idea.
The Path Forward
Start small. Pick one feature, one change, one optimization. Build the measurement infrastructure around it. Learn to love the data more than your assumptions. And remember — in a world where user behavior changes as fast as technology itself, the only sustainable competitive advantage is the ability to learn faster than your competition.
As we’ll explore in our next post on “Failure as Learning,” the teams that ship the most successful products aren’t the ones that never fail — they’re the ones that fail fast, fail cheap, and fail forward.
Now, anyone want to bet me a cup of coffee on what your next experiment should be?
What experiments are you running in your engineering organization? Share your war stories — both the wins and the spectacular failures — in the comments below.