The Most Common Way to Measure AI Visibility Is to Test as Yourself
· AI Visibility · By Chris Latham, Founder of Optimus Consulting
We were checking how AI engines described businesses while logged into our own accounts, on our own machines, with our own search history behind us. The scores looked good. They were also unrepeatable. Here is what went wrong, how we fixed it, and why it matters if anyone is selling you an AI visibility number.
We were checking how AI engines described businesses while logged into our own accounts, on our own machines, with our own history behind us. The scores looked good. They were also unrepeatable. Here is what went wrong, how we fixed it, and what to ask anyone who hands you an AI visibility number.
Here is a mistake we made, and it is worth writing up because almost everyone testing their own AI visibility is making it right now.
We were running buyer queries through the AI engines to see which businesses got named. Things like "best credit hire company UK" or "AI consultant for small business". We would type the query, read the answer, note who appeared, and score it.
We were doing that logged in. Our own accounts. Our own browsers. Our own months of history sitting behind every query.
The scores came back flattering. They were also, we eventually worked out, meaningless, which is really a question about how you measure rather than about the engines themselves.
What personalisation does to a test
The AI engines do not give everyone the same answer. They lean on what they know about you: your account, your previous conversations, your location, sometimes your browsing.
So when you ask ChatGPT or Gemini who the good providers are in your sector, and you have spent six months asking that model about your own business, you are not running a test. You are asking a system that has learned what you are interested in to tell you what you are interested in. It obliges.
The number that comes back is not wrong exactly. It is just not a measurement of anything outside your own head. Run it again next week and it moves. Run it on a colleague's laptop and it is different again. There is nothing to compare it to, including itself.
We only caught it because two people on the same query got noticeably different results and neither could reproduce the other's. That is the tell. If two people cannot get the same answer, you do not have a measurement. You have an anecdote with a number attached.
What we changed
We moved the whole thing off logged-in sessions and onto clean API calls. No account history, no personalisation, no browser. Same query, same conditions, every time.
Then we threw away the old baseline. That part was uncomfortable, because it meant admitting that some earlier readings were not worth keeping. But a baseline you cannot trust is worse than no baseline, because you will measure progress against it and believe the result.
What we run now is duller and considerably more useful:
- Clean sessions. No logged-in accounts, no history, nothing the engine can personalise against.
- Fixed queries. The exact words a buyer would use, written down and not edited between runs. Change the query and you have started a new test, not continued the old one.
- Repeat runs. Every couple of weeks, not once. A single reading tells you almost nothing.
- The same engines every time. ChatGPT, Perplexity, Google AI Overviews and Claude. Adding or dropping one mid-series breaks the comparison.
None of that is clever. It is just controlled conditions, which is the bit that gets skipped when the aim is a number for a slide rather than a number you will be held to.
Why one reading is never enough
This is not only our problem. The volatility is real and it is documented.
Analysis by Profound, looking at around 240 million ChatGPT citations, found something in the order of 40 to 60 percent of cited sources changing from one month to the next. Nate Elliott of eMarketer has made the same point qualitatively, that almost every AI search response differs from every other one, in contrast to the relative consistency of traditional Google results.
Whatever the exact figure, the direction is not seriously disputed by anyone doing this work.
That has a straightforward consequence. A one-off AI visibility snapshot cannot tell you whether you are improving. It can tell you roughly where you stand today, which is worth having, but the moment anyone draws a trend line through a single point you should stop reading.
There is a live example of why this matters this month. Google reshuffled the leadership of the division that builds Gemini in early August. Gemini powers Google's AI Overviews and AI Mode, and the flagship model has been delayed. That is a reasonable basis to expect the Google surfaces to be more volatile for a while, not less. If your numbers on those two move this quarter, some of that is Google, not you.
The llms.txt question, answered honestly
While we are being straight about measurement, here is a related one.
A lot of agencies list llms.txt as a deliverable. It is a small file you add to
your site, intended to tell AI models what your business is and how to describe it. It sounds
like exactly the sort of lever that should work.
There is no good evidence that it does. Mark Williams-Cook made the point neatly by setting up
a nonsense cats.txt file and showing that the anecdotes people cite as proof do
not demonstrate cause and effect at all. Sites that added llms.txt and then appeared more
often in AI answers were generally doing several other things at the same time.
Our position: llms.txt is cheap, it is harmless, and you may as well have one. We do not count it as a visibility lever, because we cannot show that it moves anything. If someone is charging you for it as a results-driving deliverable, ask them for the evidence and see what arrives.
We would rather say that and lose the line item than sell something we cannot stand behind.
What to ask anyone selling you a visibility number
If you are being shown AI visibility figures, by us or by anyone else, four questions will tell you how much weight they carry.
- Were these run logged in or on a clean session? If logged in, the number reflects the tester, not the market.
- What exactly was the query, word for word? Vague queries produce flattering answers. "Credit hire company Manchester" and "who should I use after a non-fault accident" are different tests with different winners.
- How many times was this run, and over what period? One reading is a snapshot. Three or more over six weeks is a measurement.
- Which engines, and did that list stay the same? Quietly dropping an engine where you score badly is the easiest way to manufacture improvement.
If the answers are good, the number means something. If they are vague, you are looking at a screenshot.
The wider point
The thing that made this worth writing up is not really about AI at all.
We audit other companies on exactly this kind of discipline. We tell them their evidence is thin, their claims are unverifiable, their numbers do not reproduce. Then we found the same fault in our own method, in the specific area we sell.
That is not a comfortable thing to publish. It is the right thing to publish, because the alternative is asserting rigour rather than showing it. Anyone can claim their audit is careful. Showing the point where it was not, and what changed, is the only version of that claim worth anything.
If you are measuring your own AI visibility, log out first. It is the single cheapest improvement available, and it will probably make your numbers worse, which is how you will know it worked.
If you would rather have someone else run it under controlled conditions, you can book a 45-minute call and we will walk through the method before you spend anything.
Related reading: Why Your Business Is Invisible to AI Search and You Are Not Buying Visibility. You Are Renting It.
Chris Latham is the founder of Optimus Consulting. He spent 25 years in motor claims, credit hire and insurance operations before moving into AI adoption and operational transformation, and now helps UK service businesses work out which parts of AI are worth their money.
Frequently Asked Questions
Why do AI engines give different answers to different people?
Because they personalise. The model draws on your account history, your previous conversations, your location and sometimes your browsing. Two people asking an identical question can get materially different answers, which is why a logged-in test cannot be compared to anything.
What is a clean session, and why does it matter?
A query run with no account, no history and no personalisation, usually through an API rather than a browser. It matters because it is repeatable. The same query under the same conditions should give a comparable answer, and comparability is the whole point of measuring.
How often should AI visibility be measured?
Fortnightly is a sensible rhythm, and never fewer than three readings before you claim a trend. Cited sources change substantially month to month, so a single reading cannot distinguish your progress from the engine's own noise.
Does llms.txt improve AI visibility?
There is no reliable evidence that it does. It is cheap and harmless to add, so there is no strong reason not to, but it should not be sold or bought as a results-driving deliverable. Ask for evidence if it is presented as one.
Which AI engines should be tested?
ChatGPT, Perplexity, Google AI Overviews and Claude cover the ground for most UK service businesses. The important thing is less which four and more that the list does not change between runs, because changing it breaks the comparison.