P-Value

Analytics

Also: Probability Value · Significance Value

P-value = probability of seeing this result, or something more extreme, if there was no real effect at all
What it measuresOdds your result is just noise
Common thresholdUnder 0.05 is treated as significant
Watch forDoesn't prove your idea is right
Depends onSample size and test duration

Quick definition

A p-value is a number between 0 and 1 that tells you how likely you would see your test results, or something more extreme, if there was no real difference between the two things you tested. Marketers use it in A/B testing to decide whether a result is worth acting on.

How it varies across Australia

Australian ecommerce and SaaS businesses typically run tests on smaller traffic volumes than larger overseas markets, so tests take longer to reach a trustworthy p-value. That favours fewer, more targeted tests over the high test-cadence approach common in bigger markets.

See conversion and testing benchmarks across Australian industries

The three numbers people mix up

Significance level(Alpha)

The threshold you decide in advance, usually 0.05, below which you'll call a result meaningful.

Set before the test
P-value(p)

The actual probability calculated after the test runs, compared against your threshold.

Calculated after the test
Statistical power(Power)

The chance your test correctly detects a real effect if one actually exists.

Depends on sample size

What it actually means

Think of a courtroom. The default assumption is innocence. The prosecution has to produce enough evidence to make that assumption look absurd. A p-value works the same way in an A/B test. The default assumption, called the null hypothesis, is that your new headline, button or price does nothing different from the old one. The p-value tells you how absurd your data makes that assumption look.

A p-value of 0.03 means that if there truly was no difference between your variants, you'd see a gap this large or larger only 3% of the time by chance. That's suspicious enough that most teams call it significant. It does not mean there's a 97% chance your variant is better. That's the single most common misreading in conversion rate optimisation (CRO).

The number is also hostage to sample size. Run a test on enough traffic and almost any tiny, meaningless difference will eventually produce a small p-value. Statistical significance and business significance are different questions, and a p-value only answers the first one.

A p-value doesn't tell you your idea worked. It tells you how surprised you should be if it didn't.

How to calculate it

P-value = probability of observing this result, or one more extreme, assuming no real effect exists

Worked example. You run an A/B test on two headlines. Headline B converts at 4.2% against Headline A's 3.8%, each on 2,000 visitors. The statistical test returns a p-value of 0.03. If there was truly no difference between the headlines, you'd see a gap this large or larger only 3% of the time by chance. Most teams treat anything under 0.05 as worth acting on.

The Australian context

Smaller Australian traffic volumes mean tests often need to run longer to reach a p-value worth trusting. A Sydney ecommerce site doing a fraction of the sessions of a comparable US site can't run the same weekly testing cadence and still get honest numbers. The fix isn't lowering your standards. It's running fewer, bigger, better-designed tests and resisting the urge to call a result early because the p-value dipped under 0.05 for a day.

Where people get this wrong

Calling a test early the moment the p-value crosses 0.05.P-values wobble as data accumulates. Checking daily and stopping at the first good-looking number inflates false positives dramatically, a problem known as peeking.
Treating a significant result as proof the change caused the lift.A low p-value means the result is unlikely to be random noise. It says nothing about whether your hypothesis for why it worked is correct.
Ignoring effect size once significance is reached.A tiny, commercially meaningless lift can still produce a low p-value on enough traffic. Statistical significance and a result worth shipping are not the same thing.

Related terms

Common questions

What p-value counts as statistically significant?

Most teams use 0.05 as the cutoff, meaning a less than 5% chance the result is random noise. That threshold is a convention, not a law. Higher-stakes decisions sometimes demand a stricter cutoff like 0.01.

Does a low p-value mean my variant is definitely better?

No. It means the observed difference is unlikely to be pure chance. It says nothing about whether the effect will hold up long term, whether your reasoning for the win is correct, or whether the lift is commercially worth the effort.

Why did my test show significance and then the result disappeared?

Usually peeking, stopping the test the moment the p-value looked good instead of waiting for the pre-planned sample size. It can also mean the effect was real but small, and later traffic simply washed it out.

How is a p-value different from a confidence interval?

A p-value is one number answering one question: is this result unlikely under no effect. A confidence interval gives a range of plausible values for the actual effect size, which is usually more useful for deciding whether a result matters commercially.

Debrief

Get the next one

No spam. No fluff. Just the next article, straight to your inbox.

Keep exploring

About New Rebellion

New Rebellion is a marketing intelligence consultancy. We build tools, score Australian businesses on how their marketing actually performs, and publish Debrief every day. This dictionary is part of how we work in the open.

How we think →