

Most AI visibility reports I see are as reliable as a single impression-share snapshot taken at 1am. Run a buyer prompt once, read your position like a SERP rank, count the mention and call it a win. But AI answers are generated afresh, not pulled from a fixed ranking. I measure them the way I learned to measure ad copy: repeated runs, a fixed prompt set and a baseline. Anything less gives a confident number that may mean nothing.
I understand why everyone starts here. You type “best project management tool for a 10-person agency” into ChatGPT, see who gets named and screenshot it for Slack. It feels like checking a SERP. I ran my first AI visibility checks the way I used to preview ads: once, at noon, then I treated what I saw as the truth.
The problem is sampling. A chatbot does not look up a stored brand ranking. It draws from a probability distribution, uses whatever it retrieves at that moment and builds an answer. Ask again and the brand list can change. When SparkToro and Gumshoe ran 2,961 prompts through 600 volunteers, fewer than 1 in 100 runs produced the same brand list; fewer than 1 in 1,000 produced it in the same order. A separate 70,000-answer study found identical brand lists on two random days only 0.3% of the time on ChatGPT and 1.1% on AI Overviews.
I used to tell clients one ad test meant nothing without a control. I was wrong to think an AI answer would be more stable. Report the hit rate, not the screenshot: run the same prompt five times and log “mentioned in 4 of 5 runs,” not “ranked second.” The first number is an observation you can repeat. The second borrows a SERP concept that does not fit.
Old SERP habits die hard. I still catch myself reading an answer top to bottom like positions one through six. Rank trackers with tidy little numbers do not help. What the deck calls AI ranking, I call a screenshot with good lighting.
The answer can change because the model generates it, sometimes using retrieved material. There is no stored order for an exact AI rank tracker to follow. Position in one answer describes that answer; it does not establish a lasting place in line. That distinction matters when someone presents a move from fourth to third as progress.
Frequency can be more useful. In the 70,000-answer test, the leading brand appeared 82% of days on ChatGPT and 89% on AI Overviews, even though the exact lists rarely repeated. I still log whether a mention appears early, middle or late. It is a diagnostic, not the headline metric. Share of runs mentioned is the number to compare over time. If presence moves from 60% across 25 runs this month to 75% next month, I pay attention. Third versus fourth in one answer gets no victory lap.
One big visibility number looks efficient on a dashboard. It also hides two different events. A mention puts your brand in the answer text; a citation links to a source. They do not necessarily arrive together. In a Writesonic analysis of roughly 16 million brand appearances, about 40% of citations did not name the source brand in the text. The reported rates were 52% on Perplexity, 49% in AI Mode, 41% in AI Overviews and 37% in ChatGPT.
A citation without a name may offer a path to your site while leaving the reader unaware of who supplied the information. A name without a citation may build awareness but give the reader no source link to follow. Neither is worthless. They are not interchangeable, either. I log two columns: named in answer and linked as source. Count mentions and citations separately before deciding what improved. Otherwise, a dashboard can show a rising total while the part you actually wanted stays flat.
A brand name in the prompt is a rigged test. Ask “is Acme Plumbing any good?” and of course Acme shows up. The question I care about is closer to “who can fix a leaking water heater in Austin this week without ripping me off?” That buyer has a problem, not a company name.
I learned this running Search. Branded terms can convert, but they tell you little about demand you have not already won. Say you spend $20k a month and 60% sits on brand. The report looks healthy while a competitor takes the unbranded queries that could grow the account. AI visibility has the same trap: score prompts that already contain your name and you have mostly measured the model’s willingness to repeat it.
Start with buyer language instead. Pull a fixed set of 10 to 15 prompts from sales calls, support tickets and the phrasing prospects use in forums. Keep branded prompts outside the scored set so they cannot flatter the result. Buyer prompts are the unbranded queries of this measurement job. If you appear in 90% of branded prompts and 10% of buyer prompts, you have a measurement hobby, not a growth signal.
One dashboard, one score, every engine averaged. I get the appeal. I once ran Search and Social under one blended CPA, which hid the fact that one channel was carrying the other for six months. A blended AI visibility score can conceal the same problem.
The surfaces do not produce the same kind of answer. One comparison found that ChatGPT named brands in 98.3% of ecommerce answers, at around six brands per answer, while Google AI Overviews named one in 7.2%. Perplexity sat elsewhere, at 84.7% with 8.49 citations. Put those outputs into one average and the number becomes easier to present, not easier to use.
Even Google’s surfaces differ. One analysis found 59% citation overlap between AI Mode and AI Overviews for top domains, versus 27% between Gemini and AI Mode and 34% between Gemini and AI Overviews. Gemini’s overlap with ChatGPT was 39%. Different retrieval and source preferences help explain why one engine may keep finding you while another does not.
Track a hit rate for each engine. A strong ChatGPT presence and a weak Overviews presence should remain two visible facts, not dissolve into a comfortable middle. The gap is what tells you where to look next.
A tool offers 500 tracked prompts and calls it coverage. I used to think a bigger keyword list meant a better Search build, back when I stacked exact-match variants like inventory. Then broad match ate the variants and taught me that volume was not structure.
Prompt wording does vary. When 600 people wrote the same headphone-buying intent in their own words, their phrasing similarity scored 0.081, while the top brands still appeared in 55% to 77% of the 994 responses. The lesson is not to invent hundreds of immaculate spreadsheet prompts. It is to represent a buying intent using the messy language people use, then hold that set still long enough to compare results.
Build a small prompt set you can run again. Here is the in-house routine I would use:

Image: "Hands Typing on Laptop Keyboard" by Image Catalog, CC0 1.0, via Flickr; modified: re-encoded as WebP and scaled to fit
The first cycle is your baseline. Score each prompt as mentioned in 0 of 5 runs, 3 of 5 or 5 of 5, then calculate hit rates by engine and by intent. Compare month two with month one using the same setup. The observed day-to-day churn in one study was 11% to 13% for core brands on the list and 65% to 78% in the tail. That is why I do not chase every small movement. My rule is to investigate an intent-level change of at least 15 points that holds across two cycles; it is a decision threshold, not a claim that everything smaller is mathematically meaningless.

Twelve real phrasings tied to buying intent, run consistently, tell me more than 500 prompts I cannot explain. No baseline, no conclusion. Until the next cycle, a good-looking first result is just a first result.
This is the one I find hardest to kill because it feels like diligence. You build the tracker, fill the sheet and watch the lines move. I have sold that feeling to clients myself, back when a thick monthly report passed for management. But watching the scoreboard never changed the score.
When you are absent from answers, the useful question is what those answers draw on instead. Source preferences differ by surface: one analysis found that authority sites supplied 26% of Gemini citations versus 10% in AI Overviews, while user content supplied 0.2% in Gemini versus 18% in Overviews. Commercial and editorial pages were the largest layer across those surfaces, at 37% to 51%. Rechecking the same prompt will not put you on a source page an engine already favors.
So I keep the monthly report short:
Ignore single-run rank, branded-prompt coverage and the blended cross-engine score. I used to hand clients a 20-tab report because thickness felt like value. It was not. It hid whether I knew which number mattered.
Then use the report to find a source gap. If ChatGPT cites you and Overviews does not, look at the pages Overviews cites for those buyer prompts. Is it a comparison page, a Reddit thread or a publisher review where you are absent? Monitoring tells you where you are missing; the source gap tells you what to work on. Address that gap, then let the next cycle show whether your hit rate changed.
That is why this myth survives. A tracker produces a tidy artifact every month, even if nobody acts on it. I measure like I measured ad copy: repeated runs, a fixed set and a baseline. But monitoring is only the receipt. The work starts when the numbers point somewhere you can actually fix.
How many times should I run the same prompt to measure AI visibility?
Run the same prompt five times and log the hit rate, for example mentioned in 4 of 5 runs. In one study of 2,961 prompts, fewer than 1 in 100 runs produced the same brand list, so a single run is not a reliable observation.
Does my brand have a stable rank in ChatGPT answers?
No. ChatGPT generates each answer fresh and there is no stored order for a rank tracker to follow, so position in one answer only describes that answer. A more useful metric is the share of runs your brand is mentioned in, compared over time.
Is a mention in an AI answer the same as a citation?
No. A mention puts your brand in the answer text, while a citation links to a source, and the two do not always arrive together. In a Writesonic analysis of about 16 million brand appearances, roughly 40% of citations did not name the source brand in the text, so count them in separate columns.
Should I include my brand name in prompts I use to track AI visibility?
No, branded prompts are a rigged test because the model will naturally repeat the name you gave it. Build a fixed set of 10 to 15 unbranded buyer prompts from sales calls, support tickets and forum language, and keep branded prompts outside the scored set.
Should I average my AI visibility score across ChatGPT, Gemini and AI Overviews?
No. The engines differ sharply, with one comparison finding ChatGPT named brands in 98.3% of ecommerce answers while Google AI Overviews did so in only 7.2%. Track a hit rate for each engine separately so a strong result on one does not hide a weak one on another.
How many prompts do I need to track AI visibility properly?
A small fixed set you can run again beats a large list. The suggested routine is 12 to 15 buyer prompts across two or three money intents, each run five times per engine on the same day, repeated monthly or biweekly with the same setup so you can compare cycles.
What should I do when my brand is missing from AI answers?
Look at what sources the answers draw on instead. Check whether the pages cited for your buyer prompts, such as comparison pages, Reddit threads or publisher reviews, are places where you are absent, then address that source gap and let the next measurement cycle show whether your hit rate changed.