An AI visibility audit measures whether AI assistants mention, cite, or recommend a brand when buyers ask them questions, and it does so with stated sample sizes, dated runs, and named engines. That last clause is the entire difference between an audit and a screenshot. This page opens up the methodology: what gets measured, why sampling is non-negotiable, and what the finished document must contain.
If you want the buyer's-eye definition and what one costs, that lives on our GEO audit page. This page is for the person who wants to know how the measurement actually works, or wants to judge whether someone else's audit is real.
The four measurements
- Mention rate. Of N samples of one buyer question, in how many was the brand named at all? This is the base metric, and it is a rate, never a rank: generated answers have no stable positions, so "you are #2 in ChatGPT" is a sentence about one moment, not a measurement.
- Share of field. Who else appears in those same answers, and how often? A brand mentioned in 3 of 24 answers means one thing when the nearest rival shows in 4, and another when the rival shows in 20.
- Citation sources. Which pages do the assistants actually cite when answering? These pages are where answers come from; the audit maps them because they are where the work happens afterward.
- The branded check. What do assistants say when asked about the brand by name? Wrong, stale, or empty branded answers are the cheapest fix and the most expensive thing to leave broken.
Why sampling is the whole game
Ask an assistant the same question twice and you will often get different brands. Cited sources turn over by roughly a third from day to day. One run is an anecdote. The discipline that turns anecdotes into measurement:
- Repetition. The same question, asked many times per engine, spread across days rather than fired in one burst, because answer variety clusters within a day.
- Disclosure. The report states n, the dates, and the engines, on the page with the numbers. A confidence interval or an honest "this sample is thin" belongs next to every rate.
- Per-engine honesty. ChatGPT, Claude, Perplexity, and Google's AI answers behave differently and serve different users. Averaging them hides exactly the differences a strategy needs. Report them separately, always.
- Kept records. The raw answers are stored, so every count can be re-derived. If an audit cannot show its answers, it is asking for faith, not offering measurement.
When you evaluate any vendor in this space, ask one question: how many samples per question, and where is that number printed? The silence you usually get back is the industry's load-bearing omission.
The anatomy of the finished document
A real audit ends in decisions, so the document is short and ordered: the question set with the evidence each question was verified against real demand; the per-question, per-engine rates with their samples; the citation map; the branded check verbatim; and a ranked fix list where every item names the question it is meant to move and the evidence it stands on. A founder should be able to read it in one sitting and know what happens next and how success will be re-measured.
What it deliberately excludes: promises of placement, position tracking dressed as science, and any number without its sample size attached.
Where this connects
Presence in AI answers is earned by the unglamorous work described in how to rank in ChatGPT: indexed pages, one canonical answer per page, legible identity, facts worth citing. The audit is the ruler for that work, before and after. Measurement without the work changes nothing; the work without measurement cannot prove it changed anything. And if you are weighing hiring any of it out, what the services cost across the market is measured on this site too.