The Debrief
L7L14L30L90All
PaidSearchIndustryTechDataBrandConversion
Tech · 7 min read13 July 2026

The Number That Moves When You Look At It

New research shows AI visibility rankings shift run to run, which means most of the tracking sold to Australian businesses reports a single snapshot as if it were a stable ranking. That is noise dressed as signal, and business owners are moving real budget because of it. Here is how to tell an estimate from a fact, and what to watch instead.

A vendor selling you both the scoreboard and the remediation is not a supplier. It is a casino that also sells you a system for beating the table.

7 min read

The Take: Most AI visibility tracking reports a single-run snapshot as if it were a stable ranking. Real measurement needs 33 to 94 repeated readings before the number holds still.

The same screenshot keeps landing in my inbox. A brand's AI visibility rank, down twelve places since last month, red arrow, panic in the caption. Someone has already booked a meeting to talk about reallocating budget. Every time, I have to be the one to say it. The number was never real to begin with.

I have worked both sides of this table. I have been the client staring at a dashboard trying to work out what to cut, and I have been the person on the agency side being paid to explain the dashboard. So let me be blunt about what is happening right now with AI visibility tracking, because a lot of Australian businesses are about to move money they should not move.

Here is the thesis, and you are allowed to disagree with it. Most AI visibility tracking sold today reports a single-run snapshot as if it were a ranking. That is static dressed up as signal. If a number changes every time you look at it, you cannot allocate capital against it. Business owners are doing exactly that anyway.

How Unstable Are AI Visibility Rankings, Really?

Across 30 platform and topic tests, AI rankings needed 33 to 94 repeated answers before they stabilised. Three tests never stabilised at all, even after 125 questions. Most tools sell you a single reading.

Start with what the tools are actually measuring. When you ask ChatGPT, Gemini or Perplexity the same question twice, you get different answers. They are built that way. Each citation is one of many URLs the model could have pulled, so the "share of voice" on your dashboard is one roll of the dice presented as a fixed fact.

New research out this month, covered by Search Engine Journal, put numbers on how much these readings move. Across 30 platform and topic tests, the number of answers needed before a ranking actually stabilised ranged from 33 to 94, counting only answers that carried citations. Three of those 30 tests never settled at all, even after 125 questions. Your dashboard pulled the data once.

A separate paper ran the same experiment under controlled conditions, identical prompts, day after day. The cited source sets overlapped just 34 to 42% between consecutive days. Roughly two thirds of the sources turned over from one day to the next (arXiv, Don't Measure Once). Same question. Same tool. A different answer most of the time.

34-42%

How much the sources cited in AI answers overlap between two consecutive days, under identical prompts. Roughly two thirds turn over. Source: arXiv, "Don't Measure Once"

This is not classic SEO. In search you move up or down a list but you stay on the list. AI search is inclusion and exclusion. You are in one answer and gone from the next. Ranking a brand on that is like scoring a cricket match by watching one over and calling the result.

They Sell You the Scoreboard and the Cure in the Same Contract

So why is the industry selling it anyway? Follow the money.

I have watched this movie three times. Rank trackers in the 2010s. Attribution dashboards in the 2020s. Now AI visibility dashboards. The plot never changes. A vendor sells you the scoreboard that tells you you have a problem, then sells you the service that fixes the problem, all in one contract. They are marking their own exam papers. The attractors will all say they are doing a really good job. They are incentivised to do that.

Even one of the researchers behind this work co-founded a company that sells the "measure it properly" version of the tool. That does not make him wrong. It means you check the maths before you sign. Rand Fishkin, whose SparkToro study found AI tools returned a different list of recommended brands more than 99% of the time you asked, gives the cleanest test I have heard. Before you spend a cent on tracking, make the provider show their maths. If they cannot, walk.

Which Businesses Are Most Likely to Buy the Static?

We measure the Australian market for a living, so measurement discipline is home turf for me. The pattern we keep seeing is uncomfortable. The businesses with the weakest data foundations are the ones most likely to buy a shiny dashboard, because they have no way to validate what it tells them.

Think about that for a second. If your own analytics are a mess, if you cannot tie a number back to revenue, then a confident vendor dashboard becomes the truth by default. You cannot argue with it because you have nothing to argue with. The tool that should be one input becomes the only input. That is how a business that cannot reliably tell you its own conversion rate ends up rearranging its content budget around a citation-share figure that will read differently tomorrow.

Across the Australian businesses we score, Data and Tracking averages 57 out of 100, the weakest of the six dimensions in our scoring methodology. It's precisely this pattern that drags the number down. Teams chase an unstable external signal instead of building their own. See how we score businesses.

The weaker your foundations, the more you outsource your judgement to whoever shows up with the most convincing screen.

Sample the innings, do not watch one over

Cricket already solved this problem. When rain interrupts a match and the conditions change, you do not just call it off or guess. The Duckworth-Lewis method extracts a fair result from incomplete, messy conditions, because someone built a system to isolate the signal from the interruption.

That is what proper measurement of a noisy channel looks like. You sample repeatedly and you read a range, not a point. The research is specific about this. For monitoring brand visibility, take at least seven runs per prompt per day. When you care about which sources get cited, at least eight (arXiv, Don't Measure Once). A single before-and-after reading cannot separate the effect of your content change from the ordinary scatter between two runs. A three-point lift in your citation share after a big content push feels like proof. It is often just the dice landing differently.

If your dashboard hands you one clean number with a decimal point on it, that decimal is not precision. It is theatre.

What I Would Actually Do About It

Treat every AI visibility number as an estimate with a range around it. Ask the vendor two questions. How many runs sit behind this figure, and what is the variance. If they cannot answer, you are not looking at measurement. You are looking at a confident guess, and you can make those yourself for free.

Watch the signals tied to money instead. Branded search volume. Direct traffic. AI-referred sessions in your own analytics. Assisted conversions. These numbers are yours and they are stable. They move only when something real happens. They will not tell you your citation share, but they will tell you whether the channel is actually putting people on your doorstep, which is the only question that pays the bills.

Then fix the boring foundations, because they are what decide whether you get cited at all. Entity clarity, so the models know who you are and what you do. Schema, so a machine can read your site cleanly. Being the primary source for something worth citing, so there is a reason to pull you into an answer in the first place. None of that shows up as a dopamine hit on a dashboard. All of it compounds.

Only after that do you pay for a scoreboard, and only one that shows its own uncertainty and tells you honestly when it does not have enough data to call a result.

The channel is real. The measurement is not there yet.

Make no mistake, this channel matters. AI is now part of daily life for 17.4 million Australians, with 5.2 million using it every day, up around 160% in a year. One in eight Australians now say an AI tool is their main way of finding information online, more than double a year ago. Few sensible operators are arguing you should ignore it.

That is exactly why the measurement grift works. The channel is real enough to feel urgent, and urgency is what makes people spend against a number they have not checked.

The winners over the next year will not be the ones with the best dashboard. They will be the ones who refuse to move money on a shifting figure while quietly building the things AI engines actually cite. Boring foundations. Real signals, not vanity ones. Estimates read as estimates, not facts. Do that, and you will still be standing when the scoreboard vendors have moved on to selling you the next one.

If your own analytics can't yet tell you which channels are real, start with a Lens scorecard and fix the foundation before you pay for another dashboard.

Frequently asked questions

Can I trust a single AI visibility ranking reading?

No. Research covering 30 platform and topic tests found rankings need between 33 and 94 repeated answers before they stabilise, and three tests never stabilised at all. A single reading is a snapshot, not a fact.

What should I track instead of AI citation rank?

Branded search volume, direct traffic, AI-referred sessions and assisted conversions in your own analytics. These are stable, tied to money and won't shift just because you asked the same question twice.

How often should AI visibility be measured to mean anything?

At least seven runs per prompt per day for general visibility, and at least eight when the question is which sources get cited. Anything less is a guess dressed up as a scoreboard.

Share this brief
Send it to a colleague who'll find it useful.
Filip Ivanković
The Debrief / From Filip Ivanković
One every morning. Six months in, you'll see the patterns most don't.
Strategy, benchmarks, and what's actually moving in Australian marketing. Four-minute read. The reps compound.
Filip Ivanković·Founder, New RebellionAboutLinkedIn