Stop A/B Testing. At Your Size It Is Statistical Theatre.
I once ran an experimentation programme that shipped more than 100 experiments a month across 10+ channels. It returned a 145% return on investment (ROI). I am the last person who should tell you to stop testing.
I am telling you to stop A/B testing.
Not forever, and not everywhere. But if you run a founder-led company with modest traffic, your A/B testing programme is almost certainly theatre. The consensus you inherited says “test everything.” That advice came from companies with millions of daily users, and it travels badly.
Here is the uncomfortable truth. A/B testing is a statistical instrument with a minimum operating scale. Below that scale, it does not degrade gracefully. It produces confident, precise-looking, wrong answers. And every week spent testing button colours is a week not spent on the bigger swings that could actually move your numbers.
This article makes four arguments. The math is against you. An underpowered test is worse than no test. Your Bayesian dashboard is not a loophole. And there is a better playbook below the threshold.
1. Run the Math Your Testing Tool Never Shows You
Every A/B test has an entry price, paid in traffic. To detect a 5% lift on a 4% baseline conversion rate, you need around 153,600 users per variant. That is over 300,000 visitors for a single test.
Hunting a bigger lift helps less than you would hope. One worked example chases a jump from 3.2% to 3.7% conversion. It still needs about 25,000 users per variant, and it assumes 8,300 eligible visitors a day to finish inside two weeks.
Now do your own arithmetic. If the page you want to test sees 10,000 visitors a month, that same test runs for five months. Seasonality, campaigns and product changes will contaminate it long before it ends.
💡Key Takeaway: If the honest duration estimate for your test is longer than a quarter, you do not have an experiment. You have a queue.
2. An Underpowered Test Is Worse Than No Test
“Some data beats no data” sounds sensible. In testing, it is backwards.
First, the peeking problem. Most teams watch the dashboard and stop the moment it turns green. Do that, and your real false-positive rate is not the 5% on the label. It is 26.1%.
Second, the winner’s curse. In a small sample, only a wildly exaggerated effect can clear the significance bar. So the wins that do appear are inflated by construction. Ronny Kohavi, who led experimentation at Bing, puts it bluntly: below 10% power, the detected effect is wrong up to half the time.
This is not hypothetical. Recast documented a geo experiment that read 11x ROI, statistically significant at the 90% level. A user-level study of the same channel put the real figure near 2.3x. Imagine the budget that moved on the first number.
No data leaves you humble. Bad data makes you certain. Certainty is what moves money in the wrong direction.
💡Key Takeaway: At low traffic, a significant result is not evidence. It is a red flag with a p-value.
3. Your Bayesian Dashboard Is Not a Loophole
The strongest counterargument I hear: “Modern tools fixed this. They are Bayesian. They work at low traffic.”
They did not fix it. A 2025 simulation ran Bayesian tests with a common stopping rule: declare a winner at 95% “probability to beat control”, checking every 100 observations. The false-positive rate hit 80%. The math is prettier. The sample is still too small.
And look at what testing achieves where it genuinely works. At Google and Bing, only 10 to 20% of experiments produce a positive result. At Microsoft, a third help, a third do nothing, and a third actively hurt. Across 20,000 experiments run by Optimizely customers, about 10% found a significant winner.
Read those numbers again. The companies that wrote the testing religion lose 80 to 90% of their bets, with millions of users and dedicated statisticians. They can afford that hit rate because each test costs them almost nothing. Each test costs you months.
A/B testing is not a rigour badge. It is a scale privilege.
💡Key Takeaway: Do not copy the ritual of companies whose constraint is not yours.
4. What to Do Below the Threshold
One practical line in the sand: under about 5,000 sessions a month on the page in question, A/B testing should not be your primary tool. Run a different playbook instead:
Ship big swings, not tweaks. Small effects are invisible at your traffic. Rewrite the whole page, change the offer, kill the weak pricing tier. Judge the change before and after against one guardrail metric.
Watch five real users. Five user-testing sessions surface around 80% of usability issues, per Nielsen Norman Group research. It is the highest-return diagnostic available at low traffic.
Ask the people who almost bought. One exit question, “What stopped you today?”, with around 50 responses, shows you the pattern.
Test demand with painted doors. Before building the feature, put up the button and count the clicks. Intent is measurable even when conversion lifts are not.
Fix the obvious for free. Slow pages, broken forms, missing trust signals. These need a bug tracker, not a control group.
Save the real A/B test for the rare decision that clears the traffic bar and genuinely divides your team.
Final Thoughts: Precision You Cannot Afford Is Theatre
The point of experimentation was never the ritual. It was better decisions per month. Below the traffic threshold, A/B testing delivers worse decisions wrapped in better-looking charts. Dropping the theatre hands you back the two scarcest resources in a founder-led company: attention and speed.
I built a career on experimentation. The programme I opened with returned 145% ROI for one reason: the volume matched the method. When your volume does not match the method, change the method. The truth does not bend to your dashboard.
If your growth engine produces dashboards instead of decisions, that is rarely a testing problem. It is a system problem, and it is fixable. Book a discovery call or connect with me on LinkedIn.
A note before you close this tab. If your top of funnel has been thinning for reasons nobody can quite explain, the cause may not sit in your strategy. It may sit in your reporting defaults. That is fixable, and the fix starts with naming what your model cannot see.
Mervyn Chua is a growth-transformation consultant helping founders and CEOs build the strategic clarity and systems to grow in an AI-first world. If this raises questions worth exploring for your brand, let’s talk.
