Why bad AI visibility advice is so hard to disprove
Most claims about what AI engines reward have never been tested against a control, and untestable claims are exactly the kind that spread. A worked example of how to read one properly, and four questions that separate a method from a story.
Ask someone working on AI visibility why they do a particular thing to a page and you will usually get a confident answer backed by real observation. Ask what happened the time they did not do it, and the conversation tends to stop.
That gap is the most interesting thing about this field right now. The problem with most AI visibility advice is not that it is wrong. Plenty of it is probably right. The problem is that it is shaped in a way that makes it impossible to check. Claims shaped that way spread faster than claims that can be tested, because nothing ever kills them.
The claim that cannot be wrong
Take any tactic in circulation. Add outbound citations to your pages. Put the year in the title. Lead every section with a direct answer. Add more structured data.
The case for each usually runs the same way: I do this on every page, and my pages perform. That is a real observation. It is also perfectly consistent with the tactic doing nothing at all. The pages someone applies their best practices to are the pages they were already investing in, on topics they already had a chance of winning.
The test that would separate those two worlds is obvious, and the habit of running it has simply not carried over to this channel. Classic search worked out split testing years ago: comparable pages, a change applied to a random half, both arms tracked while everything else moves. The AI visibility version needs the same shape, with one correction. A single page held back is not a test. It is an anecdote pointed the other way, exposed to the same drift and the same luck as any page you changed. What settles anything is a set of comparable pages, with the tactic applied to some and withheld from the rest. Both groups get measured for months, while the queries behind the prompts drift underneath them.
What makes that hard is not technical. It means arguing for pages you believe are worse, then explaining the decision to a client or a manager if the numbers dip. The professional incentive runs hard against ever finding out, and until recently the tooling to do it cheaply has barely existed.
So beliefs accumulate. Each one arrives with a plausible mechanism and a stack of supporting cases, and none has ever been isolated. Then a new retrieval layer shows up in front of search, the old certainties wobble, and the same habit generates a fresh set of rules about what language models like. The rules are new. The epistemics are not.
Why the advice is shaped this way
Two things keep the supply coming.
The first is that an unfalsifiable claim is a better product than a falsifiable one. If your pitch is that AI engines reward clarity, authority and strong brand signals, no client can ever come back and prove you wrong. If your pitch is that a specific change will move a specific measurement within six weeks, they can. Vague claims survive contact with customers. Precise ones have to earn their keep.
The second is that the message flatters its buyer. Marketing leaders have spent fifteen years with a search channel whose central currency was links from other websites. That always felt like an insult to the work. Why should someone else’s hyperlink outrank six hours of genuinely good writing? A category of advice that says the new engines reward quality, expertise and brand is telling that audience the exact thing it has wanted to hear since PageRank. You do not need to assume bad faith anywhere in the chain to get a lot of uptake. You only need a message that is pleasant to believe and expensive to verify.
The models will not referee this for you
The obvious move, when you cannot tell whether a claim is sound, is to ask an assistant. This works well for mechanism questions. Ask how retrieval augmentation works, or what a dampening factor does in PageRank, and you will get a solid answer. Those questions have settled explanations, well represented in the training data.
Ask for a strategy on a contested question and something different happens. You get the most widely repeated position, rendered fluently and with no indication of how contested it is. On AI visibility specifically, you get back the same quality and authority advice that fills the blog posts. That is what has been published most.
This closes an unhelpful loop. Volume of published opinion shapes what the models say. What the models say then gets quoted as independent confirmation. A claim can start life as marketing copy, get absorbed, and arrive back six months later carrying the authority of a neutral third party. Nobody tested it in between. On contested questions an assistant is not a referee. It is a summary of the loudest side of the argument.
A worked example: does freshness matter?
Freshness is a good specimen, because there are real numbers attached and they reward careful reading. One caution before them: visibility is not a single measurement. Being cited as a source, being mentioned by name, and being the answer the engine gives are different things, and two studies that look interchangeable may be measuring different ones.
The headline number proves nothing on its own
The widely circulated figure comes from AirOps and Kevin Indig’s 2026 State of AI Search. More than 70% of pages cited by AI have been updated within the past twelve months. That sounds decisive, and it gets quoted as proof that engines prefer recent content.
On its own it is not, because a citation rate means nothing without a base rate. People mostly ask about current things. News is a large share of general queries and is fresh by definition. Recently updated pages are also a large share of what is competitive for those queries in the first place. Fresh citations for fresh questions is roughly what you would expect in a world where recency carries no independent weight at all. The number is consistent with a freshness preference and equally consistent with there being none, which means on its own it cannot tell you which world you are in.
That is where most commentary stops, in both directions. But better data exists, in two places.
Two studies that do better
The first is Seer Interactive’s study on content recency: 7,683 pages, 47,097 citations, four months, ChatGPT, Gemini and Perplexity, four industries, with last-modified dates pulled from schema, sitemaps and HTTP headers. Its contribution is separating publish date from update date, and the split is the whole story. Measured by last update, 72% of cited content looks current. Measured by original publish date, only 42% is under a year old. More than a quarter of the pages counted as fresh were first published two or more years ago and have simply been maintained since.
The second is inside the AirOps report itself, and it is a stronger claim than the one everybody quotes. Pages that go more than three months without an update are more than three times as likely to lose visibility as recently refreshed ones. That is a longitudinal statement about what happens to the same page over time, which is much harder to explain away with base rates than a snapshot of citation ages.
Maintained content wins, not new content
Two honesty notes first. The dates in these studies are self-reported by the pages themselves, and last-modified headers are unreliable in ways that flatter freshness. Some servers re-stamp on every request, so a share of what looks recently updated is only recently re-dated. And the split needs its own base rate. Suppose most competitors refresh annually, because that is what the advice tells them to do. Then 72% maintained among the cited is close to what you would see if maintenance carried no weight at all. What the split actually rules out is the newness reading of the headline. The positive case for maintenance rests on the longitudinal number, and even that deserves a clause. Refreshed pages usually sit on sites that keep investing in other ways. The refresh is rarely the only thing that changed.
So freshness does appear to matter. But the useful conclusion is close to the opposite of what the headline number implies. It is not that engines prefer new content and you should publish more. It is that maintained content wins, and the highest-leverage work available to you is usually refreshing pages you published years ago and have not touched. “Publish and maintain” and “publish more” are very different budgets.
Notice what made that readable. A distinction between two things the headline conflated. A claim framed as a change over time rather than a snapshot. And, at the good moments, someone saying what the number would show if the effect did not exist. Those are the properties to look for.
What the engines reward, and what they verify
Underneath all of this sits a mechanical point that is easy to verify yourself.
Retrieval and citation are separate stages. A page has to be retrieved for one of the queries the engine actually runs, which leans heavily on the ordinary ranking signals for those queries. There is a stage before both, and it is easy to miss. Whether the model knows you at all shapes which queries it writes in the first place. Whether a retrieved page then gets cited is a second decision. The folk wisdom says it favours pages that answer the question directly and carry specifics: figures, dates, named comparisons, quotable sentences.
That last claim deserves the same suspicion this article has been handing out. It is plausible, widely repeated, and untested in any isolated way: folklore of exactly the shape described above. The difference is that it is cheap to hold out and check, and by now you know how.
What the system does not do is check whether those specifics are true. You can confirm this yourself, though not in the few minutes the usual telling promises, because the slow part is indexing. Invent a term, publish two lines defining it, wait until the page is being retrieved at all, then ask an assistant with browsing about it. You will often get back a confident, well-organised, entirely derived explanation, because the model’s job is to summarise what it retrieved rather than to audit it.
That cuts both ways, and both are worth holding. It tells you what a citable page probably looks like, which is genuinely actionable. It also tells you the channel rewards the appearance of specificity without checking it. So the tactics that exploit that will be common, and the countermeasures will keep moving. Any rule you learn this quarter has a shelf life. It is a good argument for measuring what engines actually say about you rather than reasoning about what they ought to say.
Four questions worth asking
The buyer’s position here is genuinely unfair: you should not have to become a search practitioner to work out whether you are being sold something real. Four questions get you most of the way without it.
What happens if we do not do this? If nobody can describe what the page looks like without the tactic, or how they would notice it failing, you are being told a story rather than shown a finding. The strongest answer anyone can give is an offer to hold the tactic back from a comparable set of pages and show you both arms.
Measured against what, over how long? A prompt run once is an anecdote. The queries engines generate behind a prompt drift week to week, so a single snapshot will cheerfully show you a win or a loss that is neither.
Can I see the work? Be wary of work located where you cannot observe it, followed by a request for trust and a retainer. Real work in this channel produces pages, mentions, rankings and answers you can go and look at.
Is this a strategy or a maintenance list? A crawl report is not a plan. Without knowing what you need to be found for, a list of broken links and overlong titles tells you what a tool noticed, not what your business needs. Most of the skill in reading one is knowing what to ignore.
The honest position
Nobody has been doing this long enough to be an expert in it, and anyone certain about the details is telling on themselves. That is not an argument for doing nothing. It is an argument for preferring measurement over doctrine. When the rules are unsettled and moving, the only durable advantage is seeing what is actually happening to you: early, and in enough detail to tell a real change from noise.
The same suspicion should extend to this article. It flatters its reader with a pleasant belief of its own: that you are the careful one, buying rigour in a market that sells comfort. Treat that the way you would treat any other message shaped to be believed.
Which is why the honest description of what Voxoria does is not that it settles anything. Your prompts run daily across ChatGPT, Perplexity, Gemini and Google AI Overviews, capturing the fan-out queries behind each one, who got mentioned, and which sources were retrieved and cited. That will not settle whether outbound citations work. Nothing settles that from the outside. What it changes is who can run the test. Apply a tactic to half of a set of comparable pages, hold the other half back, and watch both daily while the fan-out queries drift underneath them. It turns “I did this and it worked” into something you can be wrong about, which is the part the argument has been missing all along.