The question: would your brand appear again?

Your company appears in the third bullet of a ChatGPT answer. You take a screenshot, post it in Slack, and call it AI visibility. Run the same prompt again before you celebrate. The next answer may name a competitor or recommend nobody at all.

In a variance decomposition study of nearly 13,000 AI search responses, stochastic resampling within the same prompt accounted for 34.8% of outcome variance; underlying brand identity accounted for 0.7%. That does not tell you the odds of appearing in your particular query. It does tell you why a single answer makes a poor measurement.

Treat one answer like an ad with 12 clicks

I spent years watching junior media buyers pause an ad group because it logged twelve clicks and zero conversions on a Tuesday morning. A one-run AI visibility audit makes the same mistake in a newer interface. You have observed one response, not established how often a buyer will see you.

If you want to measure brand visibility across ChatGPT, Gemini, and Google AI Overviews without mistaking a screenshot for a trend, write down the question first: Across repeated runs of the same buyer prompts, how often does my brand appear, and does that rate change while an untouched control set stays steady? Then run the protocol. No prompt polishing halfway through because you dislike the first result.

Setup: lock the prompts before you run them

Build three buckets of ten

Write thirty prompts in a spreadsheet before opening an AI search engine. Keep their wording fixed on Day 1, Day 7, and Day 14. As Averi AI notes in its citation-tracking field notes, rotating a prompt set midway undermines a comparison over time. Divide the sheet into three buckets:

  1. Direct buyer prompts: High-intent comparison questions where your category is explicit and your company belongs in the consideration set, such as best autonomous ppc software for mid-market agencies.
  2. Category problem prompts: Questions about the bottleneck you solve before a buyer starts naming vendors, such as how to eliminate wasted spend in smart bidding.
  3. Untouched control prompts: Questions from a neighboring category you do not serve and will not publish content about, such as top enterprise billing automation platforms for saas.

Three specimen jars on a laboratory bench representing buyer prompts, category prompts, and an untouched control group.

The control bucket is not a place to hunt for your own brand. It tells you whether answers to untouched questions are changing too. Log which vendors appear, how long the recommendation lists are, and whether sources are cited. If vendor inclusion jumps on both your buyer questions and unrelated controls, do not rush to credit your new page. The platform may be behaving differently. The control cannot tell you why; it can stop you claiming the change as yours.

Keep each engine separate

Use the same prompts in ChatGPT with search active, Gemini, and Perplexity. Submit the corresponding queries to Google and record whether an AI Overview appears at all. Do not score a missing AI Overview as an answer that omitted your brand. Log it as no Overview, then score mentions only where one is present.

Keep your access method and conditions as consistent as you can across dates: do not mix a logged-in browser session with an API run and quietly put both in the same column. Record the engine, date, prompt, and run number. The point is not to recreate every buyer’s session. It is to avoid introducing a change you control while trying to detect one you do not.

Do not collapse the engines into one polished AI visibility score. In Profound’s analysis of citation stability across 80,000 prompts, reported monthly citation drift differed across Google AI Overviews, ChatGPT, and Perplexity. A blended percentage can hide a gain in one surface behind a loss in another. Read each engine on its own terms.

Choose 20 runs, then keep that count fixed

Run each prompt twenty times per engine at each measurement point. That is a workload, not a magic number that makes the result precise. The Don’t Measure Once evaluation addresses the problem with single-run visibility snapshots; an investigation of repeated queries across search-grounded models likewise describes citation accumulation that had not stabilized after 15 to 24 executions of the same prompt. Twenty gives you a repeatable working sample and exposes volatility a screenshot conceals. It does not make small changes decisive.

If you are doing this manually, ten runs per prompt can serve as a first pass. Label it triage, not a reliable trend. Use the same count at follow-up; changing the sample size while changing the site makes the comparison harder to read. The rule is repeat, record, and resist false precision.

What to log: four outcomes, not one checkbox

Separate a mention from a citation

For every answer, put these in separate columns:

  • Brand mentioned in body text (0 or 1): Does the answer name your company?
  • Domain cited as a source (0 or 1): Does your domain appear in references or a citation card?
  • Direct hyperlink included (0 or 1): Is there a clickable link to a page on your site?
  • Position in a recommendation list (integer or null): Are you first, third, seventh, or not in a ranked list at all?

A brand can appear in the prose while another site supplies the cited source. Averi AI’s citation-tracking notes describe why those outcomes need separate treatment. A name-drop is not a link; a link is not a top recommendation. If you collapse them into one yes/no field, you cannot tell what changed.

Cutaway diagram of an AI search response showing separate text mention, source citation, and clickable-link layers.

Record who appears beside you

Log competing vendors named in the same answer, including answers that omit you. If your mention rate falls, check whether another vendor appears more often or whether recommendation lists have simply become shorter. Those call for different next steps. One suggests a change in the comparison set; the other suggests the answer format changed. Neither is visible in a dashboard that counts only your name.

For the control prompts, apply the same discipline to vendors in that category. You are not comparing your mention rate with a billing platform’s mention rate as though they were rivals. You are checking whether the engine’s general habit of naming and citing providers moves while your site stays out of that category. A control is a weather gauge, not a second campaign.

Run the 14-day sequence without moving the goalposts

Day 1: establish the baseline

Run the twenty repetitions for every locked prompt and record every answer. Keep your planned landing-page, content, schema, or structural changes on hold for the first 48 hours while you inspect how much the baseline varies within itself. A prompt that names you in a few runs and omits you in the rest is already warning you not to overread the next percentage.

Once the baseline is logged, make the changes you planned for the buyer and category prompts. Leave the control category alone. Record what you changed and when; otherwise, you will reach Day 14 with a graph and an argument about whose homepage edit caused it.

Day 7 and Day 14: rerun the same matrix

Repeat the same prompts, run counts, and logging method on Day 7 and Day 14. Fourteen days gives you more than one follow-up view; it does not guarantee that every crawler or retrieval system has incorporated your work. An unchanged result may mean the change has not surfaced yet, not that the idea failed. A Day 2 spike, meanwhile, is not proof that your work traveled through the system overnight.

This is where the untouched bucket earns its spreadsheet space. Practitioner discussion of AI tracking tools on r/SEO reflects the frustration with dashboards that swing while nobody changes the site. If your treated prompts move and control answers also change substantially, mark the result inconclusive. Do not subtract a control-category percentage from your brand’s percentage and call the remainder causal. Different categories can move differently.

Infographic showing a 95% confidence interval for 20 runs: a 40% observed mention rate spans roughly 21.8% to 61.3%.

Read the result without pretending the band is narrow

Calculate the rate, then look at its uncertainty

Divide appearances by runs for each prompt and engine. Eight mentions in twenty runs gives an observed mention rate of 40%. It does not mean your brand has a settled 40% chance of appearing for every buyer. With a Wilson score confidence interval for a binomial proportion, that result has a roughly 21.8% to 61.3% 95% interval under the method’s sampling assumptions. The chart’s clean-looking point sits inside a wide band.

An analysis of ChatGPT Shopping offers and AEO tracking data also illustrates how rarely some product titles recur across repeated identical prompts. Take the operational lesson, not a universal benchmark: 35% on Day 1 and 45% on Day 14 is too small a change to celebrate on twenty runs alone. Check the individual prompts, the other logged outcomes, and the controls before changing a plan.

Ask whether the treated prompts separated from the controls

Look for a substantial, repeated change in buyer or category prompts that is not echoed by comparable changes in control-answer behavior. A move from four mentions in twenty runs to thirteen in twenty deserves attention that a move from four to six does not. Then inspect whether the gain persists on Day 14, whether it appears in citations and links as well as mentions, and whether controls stayed reasonably steady.

Even a sharp divergence does not prove your page edit caused it. Model updates, retrieval changes, and other events can happen during the window. The protocol gives you a better reason to investigate and a much worse excuse to declare victory early. If treated prompts and controls both swing, run the measurement again before assigning credit. A useful result survives a second look.

Separate intermittent appearances from persistent zeros

Sort the buyer and category prompts by their run histories. Put intermittent appearances in one group and prompts with no appearances in another. A five-in-twenty prompt shows that the brand can enter at least some answers under these conditions; it does not prove which page was retrieved or why the model named you. Use the Prompt-Trigger Test to investigate whether an answer draws on your page or merely includes your name alongside others.

A zero across repeated runs is a different lead. It tells you that the brand did not appear in this test, not that it has vanished from an engine’s index. Keep the prompt, engine, and dates attached to that zero. They are what make it actionable rather than an insult from a spreadsheet.

Two identical brass dice on dark slate, one showing four pips and the other two.

Turn the pattern into the next test

Intermittent mentions: inspect clarity before adding volume

If you appear in some runs but disappear in others, do not rewrite an entire page because one answer annoyed you. Check whether the page states the core answer plainly, whether important text is accessible without client-side rendering, and whether the brand and category relationship is clear. Cleaner HTML, a direct answer near the top, and explicit entity markup are candidates to test. They are not diagnoses the mention rate can make by itself.

Change the most plausible source of friction, note the change, and rerun the same prompts. If mentions become steadier while controls do not move in the same way, keep investigating that path. If nothing changes, you have learned something more useful than “write another 2,000 words”: that this edit did not show a detectable effect in the window you measured.

Persistent zeros: find the missing connection

For a prompt that stays at zero, inspect what does appear. Do the answers cite a type of page you do not have? Do they rely on third-party comparison directories, forums, or reviews where your brand is absent? A title-tag tweak cannot substitute for a page answering the buyer’s question, and a dashboard cannot create a third-party mention.

Build a loss list tying missed citations to possible content gaps. Mark whether you need a dedicated page for that query or need to examine the outside sources the engine favors. Those are hypotheses for the next cycle, not a license to manufacture citations because the score looks lonely.

Who should skip this test

Fourteen days of repeated query runs are not everyone’s best use of time. Skip this protocol for now if:

  • You spend under $10,000 a month on search and cannot trust your existing conversion tracking. If Google Ads still optimizes toward raw visits instead of qualified pipeline or closed-won revenue, fix that measurement problem before building an AI citation spreadsheet.
  • You sell an undifferentiated local service in one small area. If your immediate buyers find you through local results and directories, a large matrix of brand-comparison prompts may answer a question your customers rarely ask.
  • You only want a one-off victory screenshot. An isolated audit will age as the engines change. If you will not repeat the measurement, do not give its percentage permanent status in a report.

A visibility score built from one run per prompt is a coin flip with a chart attached. Twenty runs and an untouched control will not turn it into perfect science, but they can change your next decision. If treated prompts improve repeatedly while controls stay steady, investigate and build on the work. If both move, wait and rerun before taking credit. If neither moves, revise the content hypothesis rather than the chart.

And if running that matrix across engines becomes more mechanical labor than your team can absorb, that is the work groas is built to replace: continuous search execution with a named strategist accountable for direction and attributable revenue, not another person billing hours to admire a screenshot.