Sample Size
AnalyticsAlso: Statistical Sample · Test Sample
Quick definition
Sample size is the number of observations, visitors or responses needed in a test or survey before the results can be trusted. In marketing, it determines how long an A/B test must run, how many survey respondents you need, or how many data points a conversion rate claim requires to be statistically valid.
This estimate assumes 95% confidence and 80% statistical power. The result is visitors per variant, so double it for a two-variant test. Smaller effects and lower baseline rates require substantially more traffic.
How it varies across Australia
Most Australian businesses running A/B tests on their websites or emails are working with traffic volumes that require weeks, not days, to reach a valid sample. The mismatch between the speed marketers want answers and the sample size the data actually supports is where most testing programmes fall apart.
See conversion testing patterns across Australian industries →The four inputs that determine sample size
How often the current version converts. Lower baseline rates need larger samples.
The smallest improvement worth finding. Smaller MDE means far more traffic required.
How sure you want to be that a result isn't a false positive. Typically set at 95%.
Standard: 95%The probability of detecting a real effect when one exists. Typically set at 80%.
Standard: 80%What it actually means
Sample size is the bouncer at the door of any test or survey. Too few people in, and the results are noise dressed up as signal. The maths doesn't care how excited you are about the variant.
Think of it like flipping a coin. If you flip three times and get two heads, you wouldn't conclude the coin is biased. Flip three hundred times and the pattern starts to mean something. Sample size is that logic applied to your conversion rate test, your survey, your email subject line experiment.
In the context of A/B testing, sample size determines how many visitors each variant needs to see before you can read the result. The number depends on four things working together: your baseline conversion rate (how often the current version converts), your minimum detectable effect (how small a difference is worth finding), your desired confidence level (usually 95%, meaning you accept a 5% chance of a false positive), and your statistical power (usually 80%, meaning you want an 80% chance of detecting a real effect if one exists).
The uncomfortable truth is that the smaller the improvement you're looking for, the more traffic you need to detect it reliably. Most businesses are looking for small improvements on low-traffic pages, which means they need far longer tests than anyone wants to wait.
This interacts directly with A/B testing, conversion rate optimisation and statistical significance. Get the sample wrong and those concepts become decorative rather than useful.
The test result you stopped early to celebrate is the test result most likely to be wrong.
How to calculate it
Minimum n per variant = f(baseline rate, minimum detectable effect, confidence, power). Common approximation: n = 16 × (baseline rate × (1 - baseline rate)) ÷ (minimum detectable effect)²
Worked example. Your current checkout conversion rate is 3%. You want to detect an improvement of 0.5 percentage points (to 3.5%) with 95% confidence and 80% power. Using the approximation: n = 16 × (0.03 × 0.97) ÷ (0.005)² = 16 × 0.0291 ÷ 0.000025 = 18,624 visitors per variant. At 500 daily visitors split evenly, that's roughly 75 days per variant. Most businesses don't want to hear that number. The data doesn't negotiate.
The Australian context
Australian sites outside the major retail and financial categories often have monthly unique visitor counts in the thousands, not the tens of thousands. Running a statistically valid A/B test on a page with 2,000 monthly visitors and a 2% conversion rate can require more than six months of continuous testing. The practical answer for low-traffic Australian sites is usually to test bigger changes (higher minimum detectable effect) rather than marginal refinements, or to run user research instead of A/B tests.
Where people get this wrong
Related terms
Common questions
How do I know if my sample size is large enough?
Calculate the required sample before you start using your baseline conversion rate, the smallest improvement you care about, a 95% confidence level and 80% statistical power. When both variants have reached that number of visitors, the test is ready to read. Do not read it before then.
Can I run an A/B test on a low-traffic site?
Yes, but you need to adjust expectations. Either test bolder changes with a larger minimum detectable effect, extend the test to run for several months, or consider qualitative research methods like user interviews and session recordings instead of statistical testing.
Does sample size matter for email A/B tests?
Yes, and email platforms often mislead here. Many declare a winner after a few hundred sends, which is almost never enough to reach statistical validity on open rate differences of a few percentage points. Apply the same sample size logic you would to a website test.
What is the minimum detectable effect and how do I choose it?
The minimum detectable effect (MDE) is the smallest improvement that would be worth acting on. A practical approach: ask what conversion rate lift would meaningfully change a business decision or justify the cost of the change. Then check whether your traffic can support detecting that lift in a reasonable timeframe.
Debrief
Get the next one
No spam. No fluff. Just the next article, straight to your inbox.
Keep exploring
About New Rebellion
New Rebellion is a marketing intelligence consultancy. We build tools, score Australian businesses on how their marketing actually performs, and publish Debrief every day. This dictionary is part of how we work in the open.
How we think →