Subject Line Testing
Email MarketingAlso: Subject Line A/B Testing · Email Subject Testing
Quick definition
Subject line testing is sending two or more versions of an email subject line to small segments of a list, then sending the better-performing version to everyone else. It's a form of A/B testing focused on the single line most responsible for whether someone opens an email at all.
How it varies across Australia
Open rate as a metric has become less reliable across Australian inboxes since Apple's Mail Privacy Protection started pre-fetching images. Businesses relying purely on open rate to judge subject line tests are increasingly testing against noise. Click-through and conversion downstream matter more than they used to.
See email marketing benchmarks across Australian industries →What it actually means
Subject line testing works like a doorway. The email body is the room, but nobody sees the room if they don't walk through the door first. The subject line is the only thing doing the convincing at that moment, so testing it in isolation makes sense.
The standard version splits a small percentage of the list, sends two subject lines, waits an hour or two, then sends the winner to everyone else. This only works with enough recipients. A list of two hundred people split into two groups of a hundred each will produce results that look decisive but are mostly noise.
The bigger problem is what counts as a win. Open rate used to be the obvious metric. Since Apple's Mail Privacy Protection (MPP) started automatically opening images on delivery for iOS users, open rate is partly a measure of Apple's servers and not of human curiosity. Businesses that still optimise purely for open rate are optimising for a number that's been quietly broken.
The better approach ties the test to click-through rate or conversion rate, even if that means waiting longer for a result.
A subject line test that only measures opens is testing curiosity, not customers.
How it shows up
Subject line testing shows up as a built-in feature in most email platforms, usually as an A/B test option with a percentage split and a winner metric you choose. It shows up in reporting as two open rates sitting close together, often within a margin that isn't statistically meaningful. It also shows up in the habits of marketing teams who test every send versus teams who never test at all, both of which are usually wrong in their own way.
The Australian context
Australian email lists tend to be smaller than US or UK equivalents because the addressable market is smaller. That makes statistically sound subject line testing harder to run well. A test that would reach significance on a 500,000-person US list might never reach it on a 15,000-person Australian one.
Spam Act obligations also apply. Testing subject lines that misrepresent the email's content to boost opens (a fake urgency angle, a subject line unrelated to the body) risks both deliverability damage and consent complaints, regardless of how well the test performs.
Where people get this wrong
Related terms
Common questions
How big does my list need to be to test subject lines properly?
There's no fixed number, but as a rough guide you want at least a few thousand recipients split across test groups before the result means much. Below that, differences you see are often statistical noise rather than genuine preference.
Is open rate still a reliable metric for subject line tests?
Less reliable than it used to be. Apple's Mail Privacy Protection automatically opens images for a large share of iOS users, which inflates open rate regardless of whether a human actually engaged. Pair it with click-through rate for a truer read.
How long should a subject line test run before picking a winner?
Long enough to capture most of the opens the send will ever get, which for most lists is somewhere between two and twenty-four hours. Ending a test after ten minutes and declaring a winner is a common way to lock in a false result.
What should I actually test in a subject line?
Test the underlying angle rather than surface wording. Curiosity versus clarity, benefit-led versus urgency-led, personalised versus generic. These produce bigger and more repeatable differences than swapping a single word or emoji.
Debrief
Get the next one
No spam. No fluff. Just the next article, straight to your inbox.
Keep exploring
About New Rebellion
New Rebellion is a marketing intelligence consultancy. We build tools, score Australian businesses on how their marketing actually performs, and publish Debrief every day. This dictionary is part of how we work in the open.
How we think →