← Blog

What each AI visibility tool actually measures

We put five questions about measurement to the published documentation of seven AI visibility tools. Three collect answers in ways that are not comparable, one uses a visibility score that shares nothing but its name, and the gap between the lightest and heaviest sampling is a hundred to one.

We argued recently that asking which AI visibility tool is right is the wrong place to start. A far more useful exercise is asking five foundational methodology questions:

  1. Collection method: Does it collect answers via API, browser scraping, or clickstream, and for which engines?
  2. Sampling frequency: How many samples does it take per prompt, per day?
  3. Headline metric: What is the actual denominator of its primary visibility score?
  4. Search grounding: Are answers generated without a web search included, excluded, or blended together?
  5. Transparency: When methodology changes, is the update documented in a public changelog?

To put those questions to the test, we audited the published documentation of seven AI visibility tools (including our own) as of August 22, 2026.


Methodology Auditing Rules

To keep this comparison objective and fair, we applied three strict constraints:

  • Vendor docs only: We relied exclusively on each company’s own self-reported documentation, never competitor marketing pages or third-party comparisons.
  • Blank means missing, not hidden: A blank entry simply means our search did not yield a published answer in their public documentation; it does not explicitly mean the vendor hides it.
  • Timestamped data: Evaluated on August 22, 2026. Because these products evolve rapidly, specific technical details will inevitably shift over time.

Technical Comparison Across 7 Tools

ToolCollection MethodSampling FrequencyHeadline Visibility Score
ProfoundFront-end browser queryingOnce per day, per modelShare of total responses mentioning the brand
Peec AIDirect UI scraping (explicitly non-API)DailyResponses mentioning the brand ÷ total responses
SemrushAggregated clickstream panels + Position TrackingRolling daily; weekly brand reportsComposite index (0–100) combining topic coverage & mention consistency
EvertuneDirect API for model knowledge, plus live consumer apps100 samples per prompt, per modelShare of responses mentioning the brand (with published margin of error)
ScrunchNot foundNot foundBinary presence (explicit brand mentions only)
Achtung.appDocumented vendor APIsTwice per query, daily (Claude weekly; Lite tier 3x weekly)AI Prominence: composite (0–100) from six weighted factors
VoxoriaUI scraping + 1 direct API channelOnce per day, per engineShare of tracked answers mentioning the brand

Note: Otterly and AthenaHQ were evaluated, but neither published product-level data collection methodologies at the time of review. Otterly does publish a strong, pre-registered research methodology for its standalone studies, but not for its core tool.


4 Critical Insights Behind the Numbers

1. The Industry Uses Three Collection Models, Not Two

The debate usually gets framed as APIs vs. Web Scraping, but Semrush uses a third category: clickstream data. Instead of running scheduled automated queries, Semrush sources aggregated session data from real users.

  • Scheduled tools (Scraping/API): Let you control exact prompt sets over time, but rely on artificial querying patterns.
  • Clickstream tools: Capture true user behaviour and real-world prompt variety, but give you zero control over sample size or query frequency per prompt.

2. The Sampling Gap Reaches 100:1

Sampling rates vary drastically across vendors. Profound, Peec, and Voxoria sample once a day per prompt, and Achtung.app twice. Evertune queries each prompt 100 times per model.

On Evertune’s worked example, that changes statistical confidence dramatically:

  • 1 sample of a prompt: Carries a margin of error of roughly plus or minus 44 points.
  • 25 samples: Reduces it to about plus or minus 20 points.
  • 100 samples: Drops it to about plus or minus 10 points.

While single daily snapshots carry wide statistical noise, they accumulate value over time. In the St. Gallen convergence work, aggregating daily runs over roughly a month lands near plus or minus 7 points, and it captures model drift that a burst run at one moment cannot. Those are two separate analyses in that paper rather than a head-to-head test, so read it as the same neighbourhood of precision, not a like-for-like swap.

3. Visibility Scores Are Not Standardised

Four of the seven calculate visibility as a straight rate: (Responses Mentioning Brand ÷ Total Responses). Note the numerator. It is the count of responses that name you, not the count of mentions, so naming a brand three times in one answer earns it no more visibility than naming it once. Profound calls this explicitly binary.

Two report a composite index instead, and they are not the same composite:

  • Semrush’s AI Visibility Score (0 to 100) combines topic coverage with mention consistency.
  • Achtung.app’s AI Prominence (0 to 100) is built from six weighted factors, including mention frequency and mention breadth. Achtung notes the weighting may change as it gathers more data, though the factors themselves are stable.

So a composite cannot be compared against a rate, and the two composites cannot be compared against each other either, because they weight different things. Three of the seven headline numbers in this table are not on the same scale as the other four.

4. Two Vendors Document Search Grounding, And They Disagree

A model can answer from what it already knows or go and search first. Only two of the seven set out a rule for it, and they reach opposite conclusions.

  • Achtung.app measures search-grounded answers only, excluding closed-book replies, on the grounds that training-cutoff bias systematically favours brands that were already prominent.
  • Evertune measures both, isolating the model’s foundational knowledge through direct API access and the live consumer apps separately, then treats the gap between them as the finding: strong in training but weak in the app points to a content problem, weak in both points to a brand problem.

Both positions are defensible and they produce different numbers from the same prompt, which is the whole point. Semrush comes closest of the rest, documenting that it analyses ChatGPT “in search mode”, though it states no rule for answers that arrive without a search. The other four leave the question unaddressed in their public docs.

On changelogs, Achtung.app was the only one where we found a public record of methodology changes.


Choosing the Right Instrument

Selecting an AI visibility tool comes down to alignment with your specific analytical needs:

  • For exact end-user UI fidelity: Select tools using scraped consumer interfaces (Peec, Profound, Voxoria).
  • For single-prompt statistical precision: Choose high-frequency sample burst tools (Evertune).
  • For realistic discovery of user intent: Opt for clickstream-backed platforms (Semrush).
  • For historical metric stability: Look for vendors with public methodology changelogs.