Peeking Problem

Analytics

Also: Optional Stopping · Early Peeking

What it isChecking test results before they're ready
EffectInflates false positive rate
FixSet sample size and duration upfront
RiskCalling a winner that isn't one

Quick definition

The peeking problem is what happens when you check an A/B test's results early and stop it as soon as one variant looks like it's winning. Doing this repeatedly inflates the chance of declaring a false winner, because random noise will eventually cross a significance threshold if you keep checking.

How it varies across Australia

Testing maturity varies widely across Australian businesses. Teams running structured experimentation programs tend to pre-register sample size and duration. Teams relying on default platform dashboards are far more likely to stop tests the moment a variant looks ahead, which is exactly when the peeking problem does the most damage.

See conversion and experimentation benchmarks across Australian industries

What it actually means

Imagine flipping a coin and stopping the moment you're ahead. Flip ten times, you might be behind for a while then briefly ahead at flip seven. If you stop right there and declare the coin biased, you've fooled yourself. The peeking problem is that mistake applied to A/B testing.

Statistical significance calculations assume you decide on a sample size in advance and look at the result once, at the end. Every extra look you take before that point is another chance for random noise to cross your significance threshold purely by luck. Check a test daily for two weeks and you're not running one test, you're running fourteen chances for a false positive to sneak through.

This is why a variant can look like a clear winner on day three, then quietly lose its lead by day twelve. The early result wasn't wrong exactly, it was just noise dressed up as a signal. Conversion rate swings are naturally large on small sample sizes, and peeking rewards you for reacting to swings that would have evened out given more time.

The fix isn't complicated. Decide your sample size and test duration before launch, and don't call a winner until you hit it.

Every test you check early is a lottery ticket. Check enough times and noise will eventually look like a win.

How it shows up

It shows up as a test that gets called a winner early, ships to production, and then the lift disappears within a month. It shows up in teams that check their A/B testing tool daily and stop the moment the confidence interval looks favourable. It shows up in post-mortems where nobody can explain why last quarter's 'proven' change stopped performing. Sequential testing methods and pre-registered sample sizes are the two most common ways experienced teams avoid it.

The Australian context

Australian ecommerce and SaaS businesses running experimentation programs often work with smaller traffic volumes than their US or UK counterparts. Lower traffic means tests take longer to reach a valid sample size, which raises the temptation to peek and call things early. The instinct to move fast is understandable given tighter marketing budgets, but a false winner shipped on thin traffic tends to cost more in wasted development time than the extra weeks of patience would have.

Where people get this wrong

Stopping a test as soon as it hits statistical significance.Significance thresholds are calculated assuming one look at the end. Hitting significance on day four out of a planned four-week test is far more likely to be noise than a genuine effect.
Extending a test indefinitely, hoping it eventually shows a result.This is the same problem in reverse. Running a test past its planned sample size while checking constantly gives noise more chances to cross the threshold and produces the same false positives.
Using standard significance testing for early, frequent checks.Standard tests aren't built for repeated peeking. If a team wants to check results continuously, they need sequential testing methods designed for that, not the same maths built for a single end-of-test read.

Related terms

Common questions

Is it ever okay to look at a test before it finishes?

Looking isn't the problem, acting on it early is. Sequential testing methods are built specifically to allow safe early looks by adjusting the maths for repeated checks. Standard significance testing without that adjustment is where peeking causes damage.

How do I know if a past test result was a false positive from peeking?

Check whether the test was stopped as soon as it hit significance or ran its full pre-planned duration. If it was stopped early and the lift disappeared after launch, peeking is the likely cause. Re-running the test with a fixed sample size is the only way to confirm.

Does the peeking problem apply to multivariate tests too?

Yes, and it gets worse. More variants being tested and checked simultaneously means more chances for noise to look like a winner. Multivariate tests need larger pre-planned sample sizes and even more discipline about when results get reviewed.

What's the simplest way to avoid the peeking problem?

Calculate the required sample size before the test starts, write it down somewhere visible, and treat the test as unreadable until that number is hit. If the team can't resist checking, use a sequential testing tool built to handle repeated looks safely.

Debrief

Get the next one

No spam. No fluff. Just the next article, straight to your inbox.

Keep exploring

About New Rebellion

New Rebellion is a marketing intelligence consultancy. We build tools, score Australian businesses on how their marketing actually performs, and publish Debrief every day. This dictionary is part of how we work in the open.

How we think →