Experimentation

Why smart teams still guess wrong - the case for A/B testing

Two clever, behavioural-science-approved envelope designs lost to a plain control. A boring format change won. If experts can't predict outcomes, your roadmap can't either.

Every product roadmap is a stack of bets: this checkout flow, this pricing page, this onboarding step will move the numbers. Most digital teams have the tooling to check whether those bets pay off - and still ship most of them untested, because the team is experienced and the answer feels obvious.

Here is a great example of why experienced and obvious are not enough. It doesn't come from a SaaS product or a marketplace - it comes from an envelope.

Seven envelopes, 1.2 million households

A UK charity once tested seven envelope designs for a direct-mail fundraising campaign, sent to 1.2 million households. Before the results came in, the team ranked the designs by expected performance. The two favourites were textbook behavioural science: a salience message ("boost your donation by 25%" via gift aid) and a scarcity message ("this week only").

Both lost. Not to some other clever design - to the plain control envelope. It's the single best argument for A/B testing consulting I know, so let me show you the numbers.

The offline setting is precisely what makes the evidence so clean: one variable at a time, 1.2 million randomly assigned households, real money on the line, and the experts' predictions written down before the results arrived. Digital teams rarely get a verdict that unambiguous on their own judgment - but they place the same kind of bet every sprint.

The experiment the experts got wrong

Envelope designAvg. donation per envelopevs. control
Salience message ("boost your donation 25%")£0.18−47%
Scarcity message ("this week only")£0.28−18%
Plain control£0.34-
Portrait-format envelope£0.40+18%

The winner wasn't a message at all. It was a portrait-format envelope - the same content, rotated. It lifted average donations by roughly 18% over control, while the "smart" salience design nearly halved them.

Sit with that for a second. Professionals who design persuasion for a living, applying peer-reviewed principles, confidently backed the two worst performers out of seven. This is the honest base rate of human intuition about what works.

The translation to a digital product is one-to-one. The envelope is your push notification, subject line or landing page above the fold. The salience and scarcity messages are the copy variants sitting in your next growth sprint. The portrait format is the kind of layout change a stakeholder waves through as cosmetic - and it was the only thing that worked. Different channel, identical decision: several plausible options, strong opinions about which will win, and no way to know without measuring.

If experts can't predict which envelope wins, nobody in your roadmap meeting can predict which feature wins. That's not cynicism - it's the entire business case for experimentation.

Industry-wide, the numbers back this up: most mature experimentation programs report that only somewhere between one in five and one in three tests produce a clear win. The value of a testing practice isn't that it makes you right more often. It's that it stops you from shipping the £0.18 envelope at full scale.

What a working experimentation practice looks like

I've helped marketplaces, e-commerce companies, app products and SaaS teams build or repair their testing practices. The tooling is rarely the problem - Optimizely, GrowthBook, an in-house assignment service, all fine. The practice around the tooling is where programs live or die. Five things matter most.

1. Hypotheses worth testing

A hypothesis is not "we think the new checkout is better." It's a falsifiable sentence with a mechanism: "Showing the delivery date on the product page will increase checkout completion by reducing uncertainty, measured as +2% or more on order conversion."

2. Power and sample size, decided before launch

The most common failure I find in audits: tests that never had a chance. With a 4% baseline conversion and traffic for 20,000 users per variant, you cannot reliably detect a 5% relative lift. You'll run for two weeks, see nothing significant, and wrongly conclude the idea failed.

3. Guardrail metrics

Every test needs a success metric and two or three guardrails - metrics the variant must not damage. A checkout test that lifts conversion 3% while quietly increasing refund rates 8% is a loss wearing a winner's medal. Short-term conversion wins that erode repeat behaviour are especially sneaky; that's exactly the blind spot that customer lifetime value prediction exists to close, and why we often wire predicted LTV in as a guardrail.

Guardrails only work if metric definitions are agreed before the test. If your teams still argue about what "conversion" means, fix your KPI framework first - experimentation on top of contested metrics just industrialises the arguing.

4. No peeking

Checking results daily and stopping the moment the dashboard turns green roughly triples your false-positive rate. A test that "reached significance" on day 4 of a planned 21 is not a winner; it's a coin that happened to land heads a few times in a row while you were watching.

Two honest options: commit to the pre-computed sample size and only decide at the end, or use a sequential method (like mSPRT or group-sequential boundaries) that is built for continuous monitoring. What you can't do is use fixed-horizon statistics with sequential behaviour.

5. Decision memos

Every concluded test gets a one-page memo: hypothesis, design, result with confidence interval, guardrail readout, decision, and what we'd test next. It sounds bureaucratic. It's the opposite - it's the institutional memory that stops teams from re-running 2023's losing test in 2026 because everyone who remembered it has left.

A case from the field

An e-commerce scale-up brought us in because their experimentation program "wasn't finding wins anymore." The audit told a familiar story: around 60% of their tests were underpowered from day one, results were peeked at daily, and there was no record of past experiments beyond screenshots in Slack.

We didn't add tooling. We added discipline: a pre-registration doc with power calculation required to launch, three company-level guardrails (margin per order, 60-day repeat rate, support-ticket rate), a no-peeking rule with a sequential option for urgent tests, and a decision-memo archive.

Two quarters later the numbers moved in the right direction: the share of tests ending in a clear, documented decision went from roughly a third to over 80%, and the team shipped two unglamorous winners - a shipping-cost display change and a shorter registration form - worth an estimated 2–3% on order conversion combined. Fewer tests, more answers.

Common failure modes to watch for

The takeaway

You don't run experiments because you distrust your team's ideas. You run them because the envelope data - and twenty years of industry results - say that nobody's ideas are reliably right, including the experts'. The companies that win aren't the ones with the best intuition; they're the ones with the cheapest, fastest, most honest way of finding out.

Test the boring thing. Trust the process, and compound growth.

Want experiments that end in decisions?

We audit and build A/B testing practices for marketplaces, e-commerce, fintech, SaaS and consumer apps - power discipline, guardrails, and a decision culture that survives the HiPPO. Tell us where your testing program hurts.

Start a conversation