MDE and Sample Size for CRO

Conversion & UX

Also: Minimum Detectable Effect · Sample Size Calculation · Statistical Power for A/B Testing

Sample size rises fast as your Minimum Detectable Effect (MDE) shrinks. Small MDE needs a lot more visitors.
MDE meansSmallest lift worth detecting
Smaller MDENeeds far more traffic
Common mistakeEnding tests before power is reached
Runs longLow-traffic sites, tiny expected lifts

Quick definition

Minimum Detectable Effect (MDE) is the smallest improvement in conversion rate a test is designed to reliably detect. Sample size is the number of visitors needed to detect that effect with confidence. The smaller the lift you want to catch, the more traffic the test needs before you can trust the result.

Run the numbers
%
%
Approximate visitors needed per variant12,933

This is a simplified approximation for planning purposes. Use a proper sample size calculator before committing a test budget.

How it varies across Australia

Most Australian sites running conversion rate optimisation (CRO) programs underestimate the traffic a test actually needs. Sites with modest monthly traffic typically need to test for large, obvious changes rather than small tweaks, or accept much longer test durations.

See conversion benchmarks across Australian industries

The two numbers that decide if your test can work

Minimum Detectable Effect(MDE)

The smallest lift in conversion rate you've decided is worth catching. Set before the test starts, not after.

Smaller MDE, more traffic needed
Sample size

The number of visitors per variant required to detect the MDE at your chosen confidence level.

Calculated, not guessed
Statistical power

The probability your test detects a real effect if one exists. Usually set at 80 percent.

Underpowered tests miss real wins

What it actually means

Think of Minimum Detectable Effect (MDE) as the size fish you're fishing for. If you're trying to catch a whale, a small net works fine because whales are hard to miss. If you're trying to catch a minnow, you need a much finer net and a lot more patience. A/B testing works the same way. Detecting a large lift in conversion rate needs relatively little traffic. Detecting a small lift needs a lot more.

Most teams skip this step entirely. They launch a test, watch the dashboard, and call a winner the moment one variant looks ahead. That's not statistics, that's pattern matching on noise. Without setting an MDE and calculating sample size upfront, you don't know if you're looking at a real signal or random variation in conversion rate.

The practical effect is that low-traffic sites need to test big swings, headline rewrites, pricing changes, layout overhauls, rather than small tweaks like button colour. A site with modest daily visitors chasing a small MDE on a low-traffic page can run a test for months and still not reach a reliable answer. That's not a testing failure. That's the maths doing its job and telling you the test was underpowered from the start.

A test with no sample size calculation isn't an experiment. It's a guess with a dashboard attached.

How to calculate it

Sample size depends on baseline conversion rate, the MDE you want to detect, and your chosen confidence and power levels.

Worked example. A page converts at 3 percent. You want to detect a lift to 3.6 percent, a 20 percent relative improvement. At 95 percent confidence and 80 percent power, that test needs roughly 14,000 visitors per variant. Halve the MDE to a 10 percent relative lift and the required sample size roughly quadruples.

The Australian context

Australian sites outside the largest ecommerce and finance brands often run at traffic volumes where small MDE tests simply aren't viable in a reasonable timeframe. A site with a few thousand monthly visitors chasing a small lift on a mid-funnel page may need several months to reach a reliable sample. The honest move in that situation is testing bigger changes less often, rather than running many small tests that never reach statistical power.

Where people get this wrong

Calling a test early because one variant is ahead.Early leads in a test are common and often reverse. Without reaching the calculated sample size, an early lead is noise, not a result.
Setting the MDE after seeing the data.Picking the smallest lift that happens to look significant, after the test has run, is a way of fooling yourself. The MDE has to be set before the test starts.
Assuming more traffic always solves the problem.Traffic helps, but if the true effect is smaller than your MDE, no amount of traffic will make the test conclusive. Sometimes the answer is testing a bigger change instead.

Related terms

Common questions

What is a good Minimum Detectable Effect for a CRO test?

There's no universal good number. It depends on your traffic and how long you can afford to run the test. Higher-traffic sites can afford to detect smaller lifts. Lower-traffic sites need to target larger, more obvious changes to reach a conclusive result in a reasonable time.

Why did my A/B test show a winner that didn't hold up later?

This usually happens when a test is stopped before reaching its calculated sample size. Early results fluctuate a lot. Calling a winner based on an early lead is one of the most common causes of false positives in conversion rate optimisation.

How long should I run an A/B test?

Long enough to reach the sample size your MDE calculation requires, not a fixed number of weeks. Traffic volume, baseline conversion rate and the size of the effect you're trying to detect all change the required duration.

Can I test small changes on a low-traffic site?

You can, but the test will likely run for a long time or fail to reach statistical power. On low-traffic sites it's usually more productive to test bigger, bolder changes that produce a larger effect and reach a reliable answer faster.

Debrief

Get the next one

No spam. No fluff. Just the next article, straight to your inbox.

Keep exploring

About New Rebellion

New Rebellion is a marketing intelligence consultancy. We build tools, score Australian businesses on how their marketing actually performs, and publish Debrief every day. This dictionary is part of how we work in the open.

How we think →