Statistical Power
AnalyticsAlso: Power · Test Power
Quick definition
Statistical power is the probability that a test correctly detects a real effect when one genuinely exists. Low power means a test can run cleanly and still miss a true winner, reporting no difference when there actually was one. It's calculated from sample size, effect size and the significance threshold you've chosen.
How it varies across Australia
Australian ecommerce and lead-gen A/B testing programmes tend to run underpowered more often than overpowered, mostly because traffic is smaller than teams assume. Tests get called early, and small businesses lose more real winners to missed detection than they lose to false positives.
See conversion testing patterns across Australian industries →What it actually means
Statistical power is the answer to a question most A/B testing dashboards never ask out loud: if there really is a difference between your two versions, how likely is your test to actually find it?
Think of it like a metal detector at the beach. A powerful detector picks up a coin buried under sand. A weak one walks right over it and reports nothing there. The coin was real. The detector just wasn't sensitive enough to catch it.
Power depends on three things working together. Sample size, the size of the effect you're trying to detect, and the significance threshold you've set for calling a result real. Small sample, small effect, strict threshold. All three push power down.
Most conversations about A/B testing focus on statistical significance, the p-value that says a result probably isn't noise. Power is the quieter, more important twin. A significant result with adequate power is trustworthy. A non-significant result from an underpowered test tells you almost nothing, because the test was never likely to detect the effect even if it was there. That's the trap. Teams treat 'no significant difference' as 'no difference' when it often just means 'we didn't run it long enough to see.'
An underpowered test doesn't fail loudly. It just quietly tells you nothing happened when something did.
How to calculate it
Power depends on sample size, minimum detectable effect and significance level, calculated together rather than read off a single formula
Worked example. You're testing a new call-to-action expecting a lift of a few percentage points in conversion rate. With your current traffic, a power calculator tells you that you'd need several weeks of data to reliably detect that size of lift. If you call the test after four days because the numbers look promising, you haven't reached adequate power. Whatever the dashboard shows, you can't trust it yet.
The Australian context
Australia's smaller market means most businesses running conversion-rate optimisation (CRO) tests have less traffic than their US or UK counterparts running the same playbook. A test design that works fine for a high-traffic US retailer can be badly underpowered for an equivalent Australian site. Local teams often need to test bigger changes, run tests longer, or pool traffic across more pages to reach adequate power. Copying overseas testing cadence without adjusting for smaller sample size is one of the more common quiet failures in Australian CRO programmes.
Where people get this wrong
Related terms
Common questions
What is a good statistical power for an A/B test?
A conventional minimum many teams use is that a test should be more likely than not to catch the effect, with common practice aiming well above a coin-flip chance. The right level depends on how costly a missed real effect would be to your business, not on a fixed rule.
How do I increase the power of my test?
Run the test longer to gather more sample, test a bigger change so the effect size is larger and easier to detect, or reduce the number of variants competing for traffic. Increasing traffic to the test, where possible, is usually the most reliable lever.
Is low power the same as a false positive?
No. Low power relates to false negatives, missing a real effect. A false positive is calling a result real when it's actually noise. They're opposite failure modes and both are shaped by how the test was designed.
Can a test be significant but still underpowered?
Yes, though it's less common. A significant result from an underpowered test is more likely to be an inflated or unstable estimate of the true effect. Power calculations matter most for interpreting non-significant results, but they affect confidence in significant ones too.
Debrief
Get the next one
No spam. No fluff. Just the next article, straight to your inbox.
Keep exploring
About New Rebellion
New Rebellion is a marketing intelligence consultancy. We build tools, score Australian businesses on how their marketing actually performs, and publish Debrief every day. This dictionary is part of how we work in the open.
How we think →