Statistical Power

Analytics

Also: Power · Test Power

Power = probability a test detects a real effect if one exists
What it measuresChance of catching a real effect
Low power riskReal wins get missed
Depends onSample size and effect size
Common thresholdA conventional minimum, not a law

Quick definition

Statistical power is the probability that a test correctly detects a real effect when one genuinely exists. Low power means a test can run cleanly and still miss a true winner, reporting no difference when there actually was one. It's calculated from sample size, effect size and the significance threshold you've chosen.

How it varies across Australia

Australian ecommerce and lead-gen A/B testing programmes tend to run underpowered more often than overpowered, mostly because traffic is smaller than teams assume. Tests get called early, and small businesses lose more real winners to missed detection than they lose to false positives.

See conversion testing patterns across Australian industries

What it actually means

Statistical power is the answer to a question most A/B testing dashboards never ask out loud: if there really is a difference between your two versions, how likely is your test to actually find it?

Think of it like a metal detector at the beach. A powerful detector picks up a coin buried under sand. A weak one walks right over it and reports nothing there. The coin was real. The detector just wasn't sensitive enough to catch it.

Power depends on three things working together. Sample size, the size of the effect you're trying to detect, and the significance threshold you've set for calling a result real. Small sample, small effect, strict threshold. All three push power down.

Most conversations about A/B testing focus on statistical significance, the p-value that says a result probably isn't noise. Power is the quieter, more important twin. A significant result with adequate power is trustworthy. A non-significant result from an underpowered test tells you almost nothing, because the test was never likely to detect the effect even if it was there. That's the trap. Teams treat 'no significant difference' as 'no difference' when it often just means 'we didn't run it long enough to see.'

An underpowered test doesn't fail loudly. It just quietly tells you nothing happened when something did.

How to calculate it

Power depends on sample size, minimum detectable effect and significance level, calculated together rather than read off a single formula

Worked example. You're testing a new call-to-action expecting a lift of a few percentage points in conversion rate. With your current traffic, a power calculator tells you that you'd need several weeks of data to reliably detect that size of lift. If you call the test after four days because the numbers look promising, you haven't reached adequate power. Whatever the dashboard shows, you can't trust it yet.

The Australian context

Australia's smaller market means most businesses running conversion-rate optimisation (CRO) tests have less traffic than their US or UK counterparts running the same playbook. A test design that works fine for a high-traffic US retailer can be badly underpowered for an equivalent Australian site. Local teams often need to test bigger changes, run tests longer, or pool traffic across more pages to reach adequate power. Copying overseas testing cadence without adjusting for smaller sample size is one of the more common quiet failures in Australian CRO programmes.

Where people get this wrong

Calling a test early because it hit significance.Peeking at results before reaching planned sample size inflates false positives and undermines the power calculation the test was designed around.
Treating 'not significant' as proof there's no difference.An underpowered test can fail to detect a real effect. Absence of a significant result is not evidence of absence, it's often just evidence of insufficient power.
Never calculating required sample size before starting.Without a pre-test power calculation, there's no way to know if the test could ever have detected the effect size you cared about, win or lose.

Related terms

Common questions

What is a good statistical power for an A/B test?

A conventional minimum many teams use is that a test should be more likely than not to catch the effect, with common practice aiming well above a coin-flip chance. The right level depends on how costly a missed real effect would be to your business, not on a fixed rule.

How do I increase the power of my test?

Run the test longer to gather more sample, test a bigger change so the effect size is larger and easier to detect, or reduce the number of variants competing for traffic. Increasing traffic to the test, where possible, is usually the most reliable lever.

Is low power the same as a false positive?

No. Low power relates to false negatives, missing a real effect. A false positive is calling a result real when it's actually noise. They're opposite failure modes and both are shaped by how the test was designed.

Can a test be significant but still underpowered?

Yes, though it's less common. A significant result from an underpowered test is more likely to be an inflated or unstable estimate of the true effect. Power calculations matter most for interpreting non-significant results, but they affect confidence in significant ones too.

Debrief

Get the next one

No spam. No fluff. Just the next article, straight to your inbox.

Keep exploring

About New Rebellion

New Rebellion is a marketing intelligence consultancy. We build tools, score Australian businesses on how their marketing actually performs, and publish Debrief every day. This dictionary is part of how we work in the open.

How we think →