One ChatGPT screenshot is not an AI visibility report. It is one answer to one prompt at one moment, about as useful as judging a Google Ads campaign after four clicks. An AI engine can cite your brand on one run and omit it on the next without a single change to your website.

If you want to know which buyer questions trigger a citation—and whether your content earned it—run a test. Keep the prompts fixed, repeat them, and leave comparable pages untouched as controls. A before-and-after comparison without a concurrent control group cannot separate your edit from changes to an engine’s index or a competitor’s site. I would not accept a cropped screenshot as proof of a bidding win. I will not accept one here.

Question: Did the page change earn more citations?

The test asks whether one structural or editorial change to a set of pages increases citations for those pages in answers to qualified buyer questions. The primary comparison is the change in citation rate for treated pages minus the change for untouched control pages. Brand mentions, recommendation position, and competitor citations are worth logging, but they do not replace that comparison.

My expectation: a change that makes a page easier to retrieve and extract may improve its citation rate. The control group is there to challenge that expectation. If treated and control pages move together, I have no reason to credit the edit, however good the post-test screenshot looks.

Why the same prompt needs more than one run

Hold the question still; let the answers vary

Identical queries can produce different citations across runs. Phrasing introduces another variable: add a constraint such as “for mid-market teams” or “under $500/month,” and you may get a different answer and a different set of sources. Change the wording halfway through a test and you have changed the test.

The same query enters three paths and produces different sets of brand citations.

Engines differ, too. ChatGPT Search and Perplexity have shown limited citation overlap on identical queries. A citation in one is not evidence of a citation in the other. Track each engine separately before you look at a combined result.

Practical rule: one run can give you an example to inspect. It cannot establish a citation rate.

Setup: Lock the prompts, pages, and measurement

Build 30 prompts from buyer language

Do this before editing any page. Pull candidate questions from sales calls, win/loss notes, and buyer interviews rather than pasting an old keyword list into a chat box. The prompt set should reflect how prospects ask for help, including the constraints they bring to a purchase.

Use 30 prompts, divided into three intent tiers:

  1. Category discovery — 10 prompts. The buyer knows the problem but not the tools: What software helps ecommerce brands automate inventory reordering without manual spreadsheets?
  2. Comparative evaluation — 10 prompts. The buyer is weighing named options: Compare [Competitor A] and alternatives for mid-market B2B lead scoring.
  3. Constraint-heavy implementation — 10 prompts. The buyer specifies a budget or operating model: What paid search management options exist with flat monthly pricing rather than a percentage of ad spend?

Assign each prompt to the page and topic it is meant to test. Keep the wording and that assignment fixed for the full experiment. If a buyer’s phrasing gives you a better idea on day four, save it for the next test. Do not sneak it into this one.

Choose the answer engines

Run every prompt in ChatGPT Search, Perplexity AI, and Google Gemini or Google AI Overviews. Record which third surface you chose and keep it the same throughout. This is a three-engine test, not an excuse to treat their answers as interchangeable.

Use dedicated scripts or fresh, private browser sessions with memory disabled. Do not run the baseline from your daily account and the follow-up from a clean session. If an answer draws on earlier chat context, discard that run and repeat it under the agreed conditions.

Split ten comparable pages into two cohorts

Select ten pages covering distinct but commercially comparable topics. Randomly assign five to the treated cohort and five to the control cohort. Pair each page with prompts relevant to its topic so both cohorts get a comparable mix of the three intent tiers. Keep that page-to-prompt mapping for the baseline and the follow-up.

Treated and control web pages shown as two parallel tracks in a split test.

An untouched control group is not a group you edit less enthusiastically. Leave its copy, schema, and links alone during the test window. Document any site-wide change you cannot avoid; it may affect both cohorts. If treated and control citations rise together, that is exactly the possibility the controls are meant to reveal.

Baseline week: Find out how noisy the answers are

Run each prompt five times per engine

Spend seven days observing before you change the treated pages. Run each of the 30 prompts five times in each of the three engines, spacing runs across days and hours. That produces 450 baseline evaluations, with 15 prompt-and-engine runs for each prompt across the week.

Five repeats are a practical starting point, not a guarantee of certainty. A prompt that earns two citations in five runs has a measured rate of 40% in that small sample. A single run would have labelled it either 0% or 100%. Neither label deserves a slide deck.

Give every run its own row

Record the output, not your impression of whether the answer was “favourable.” I have spent enough time in search-term reports to know how quickly a tidy summary can hide the thing you needed to inspect. Capture:

  • Brand mention (1/0): Does the generated answer name your company?
  • Citation to your domain (1/0): Does a clickable source link point to your site?
  • Citation to the assigned target page (1/0): Does that link point to the page paired with this prompt? This is the primary page-level measure.
  • Recommendation rank: If the answer lists vendors, where does your brand appear?
  • Cited URL and competitors: Which pages did the engine actually cite?

A mention is not a citation. A homepage citation is not a citation to the treated landing page. Keeping those distinctions in the sheet prevents a brand-level improvement from being mistaken for proof that your edit worked.

Calculate the baseline target-page citation rate separately for treated and control cohorts: target-page citations divided by runs assigned to each cohort. Keep engine-level rates visible as well. Record how much results vary across repeats; you will need that context when the follow-up produces an attractive number.

Intervention: Change one kind of thing

Pick one change type for the five treated pages. Do not rewrite the copy, rebuild the page delivery, and revise the schema in the same window. I would not change an ad headline and its destination URL together, then pretend I knew which one moved CPA. The same discipline applies here.

Research into ChatGPT citations has examined features including Q&A structure and direct-answer formatting. If you choose an editorial intervention, place a concise, approximately 50-word answer to the primary buyer question near the top of each treated page. Remove the corporate throat-clearing it replaces. Leave the page’s technical structure alone.

If you choose a structural intervention, keep the prose unchanged and make the same kind of delivery or markup fix across the five pages. The draft options are reducing client-side JavaScript that obstructs access to the text or cleaning up JSON-LD schema. Pick one. Then leave the control pages strictly untouched.

Write down what changed and when it went live. Otherwise the test becomes a memory exercise with a spreadsheet attached.

Follow-up: Repeat the run, not just the screenshot

Do not check ChatGPT an hour after publishing and call the answer your result. Allow 14 to 21 calendar days before the follow-up runs. During that interval, check server logs for visits to the modified URLs from relevant crawlers, such as OAI-SearchBot, Bingbot, or PerplexityBot. A crawler visit does not promise a citation; a missing visit gives you a useful reason to investigate before interpreting silence as rejection of the copy.

Then repeat the baseline sequence: the same 30 prompts, the same three engines, five runs per prompt per engine, spread over seven days. That is another 450 evaluations. Keep prompt wording, page assignments, session conditions, and logging fields unchanged.

The unglamorous part is the test. If the post-test sheet is less disciplined than the baseline sheet, the comparison is not worth the time you spent building it.

Reading the result: Subtract what happened without your edit

Use a difference-in-differences calculation for the primary target-page citation rate:

Net lift = (Treated post − Treated baseline) − (Control post − Control baseline)

Suppose the treated cohort rises from 8% to 16%, while the untouched control cohort rises from 10% to 18%. Each gained eight percentage points. The net lift is (16% − 8%) − (18% − 10%) = 0 percentage points. A before-and-after slide would call the treated result a win. The control group says otherwise.

Three possible test patterns for AI brand citations: treated-page lift, shared movement, and competitor movement.

Do not promote a small positive difference into a finding just because the formula returns a number above zero. Compare it with the variation you saw across repeated runs. Repeated-run analysis can use uncertainty intervals to assess whether a shift is credible. With this protocol’s limited sample, a result that sits inside the observed noise calls for another test, not a victory lap. Even a clear separation supports a narrower conclusion than “we own AI search”: this intervention improved measured citations for these prompts, pages, engines, and dates.

Check what moved when your page did not

A third pattern deserves a look. Your treated pages may earn no more citations while a different competitor starts appearing in answers to the same comparison prompts. Log the competing URL rather than filing the run under “no change.” It may show you which page the engine used to answer the buyer’s question.

Inspect that page for things you can actually compare: a table, pricing information, or a concise specification list. Do not assume you know why the engine chose it. Use the observation to choose the next intervention. A competitor citation is evidence about the answer, not proof that your own page was nearly chosen.

Decide the next test from the outcome

  • If treated pages pull ahead of controls beyond the observed noise: Keep the change documented. Test the same intervention on additional pages before treating it as a rule for the whole site. The result gives you a promising playbook, not a lifetime guarantee.
  • If treated and control pages move together: Do not credit the edit. If the intervention was editorial, check whether crawlers reached the changed pages and whether their text was accessible before rewriting more copy. If access looks sound, the next test needs a different hypothesis.
  • If competitor citations move while yours do not: Inspect the cited URLs and design a focused follow-up. If those prompts matter commercially, consider targeted paid search while you work on organic citations. Paid coverage is not evidence that the page fix succeeded; it is a separate response to missed demand.

That is the decision this test exists to support. Not “are we visible?” but what would we change next, and what evidence earns that change?

Automate the repetition, not the conclusion

You can run the first cycle in a spreadsheet. You probably will not want to do it by hand every month. Prompt-tracking tools such as Profound, Peec AI, and Otterly can handle much of the repeated querying and mention tracking. A monitoring dashboard still does not edit a page or decide which missed buyer question is worth pursuing.

That is the distinction behind groas: it pairs autonomous execution across organic and paid search with a named human strategist responsible for direction and outcomes. The point is not a prettier chart of missing citations. It is to spend less human time collecting repetitive observations and more on the commercial decision they support.

Copyable run log

Give every evaluation one row. This compact template keeps the primary measure beside the supporting observations; add your exact prompt text and full URLs in the working sheet.

| Run ID | Date & time | Engine | Exact prompt | Cohort | Target page | Brand mentioned? | Domain cited? | Target page cited? | Rank | Cited URL / competitors | | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | | 001 | Baseline, day 1 | ChatGPT Search | Compare [Brand] vs [Competitor] for mid-market inventory automation | Treated | /solutions/inventory | 1 | 1 | 1 | 1 | Target page cited | | 002 | Baseline, day 1 | Perplexity AI | What inventory tools integrate with Shopify Plus under $500/mo? | Control | /pricing | 0 | 0 | 0 | N/A | Competitor pricing page cited |

Keep the controls untouched. Log the repeats. If the treated pages do not beat the control movement, do not sell yourself the screenshot. In search marketing, whether paid or earned, the math still gets the last word.

Frequently asked questions

Why can't I just take one screenshot of ChatGPT mentioning my brand as proof of AI visibility?

A single screenshot is one answer to one prompt at one moment. AI engines can cite a brand on one run and omit it on the next without any change to the website, so one run can give you an example to inspect but cannot establish a citation rate.

What is the control group for in an AI citation test?

Untouched control pages separate the effect of your edit from changes to an engine's index or a competitor's site. If treated and control pages move together, the edit does not get credit, however good the post-test screenshot looks.

How many prompts should I use in an AI citation test and where do they come from?

Use 30 prompts divided into three intent tiers: 10 category discovery, 10 comparative evaluation, and 10 constraint-heavy implementation prompts. Pull candidates from sales calls, win/loss notes, and buyer interviews rather than old keyword lists, and keep the wording fixed for the full experiment.

Do I need to test multiple AI engines, or is ChatGPT enough?

Run every prompt in ChatGPT Search, Perplexity AI, and Google Gemini or Google AI Overviews, and track each engine separately. ChatGPT Search and Perplexity have shown limited citation overlap on identical queries, so a citation in one engine is not evidence of a citation in another.

How many times should I repeat each prompt when measuring AI citations?

Run each of the 30 prompts five times in each of the three engines, spaced across days and hours over a seven-day week, which produces 450 baseline evaluations. Five repeats are a practical starting point; a single run would label a page either 0% or 100% cited.

Can I change the copy, schema, and page delivery at the same time when testing for AI citations?

No. Pick one change type for the treated pages, such as adding a concise roughly 50-word answer near the top, reducing client-side JavaScript, or cleaning up JSON-LD schema. Changing several things at once makes it impossible to know which one moved the citation rate.

How do I know whether a change to my page actually earned more AI citations?

Wait 14 to 21 days, then repeat the full baseline sequence with the same prompts, engines, and run counts. Calculate the net lift as the change in citation rate for treated pages minus the change for untouched control pages, and compare any positive result against the variation observed across repeated runs.