A screenshot of missing AI citations is not evidence that buying new pages will get you cited. I get versions of that screenshot twice a month: buyer prompts in a grid, a competitor in every answer, my reader’s brand nowhere. Underneath sits the proposal to close the gaps. The missing part is a test showing that closing one changes the answer.
Most platforms will sell you the gap list and stop there. Treat the list as a set of hypotheses. Split the prompts in half, fix one half, hold the other back, and watch citation rate for six weeks. If the fixed half does not beat the holdout by more than normal movement, you did not buy visibility. You bought pages.
Question: does a closed gap change who gets cited?
A gap report can mean two different things. Open-source GEO trackers call a prompt where a competitor is mentioned and you are not a Visibility Gap. A Citation Gap is stricter: a competitor domain is cited and yours is not (tracker definitions). For this test, use citation gaps. A mention can tell you where to investigate, but it cannot answer whether a new passage earns a citation.
My expectation: a fix can raise citation rate when a retrievable passage answers the buyer’s question better than the page previously cited. That is the mechanism to test, not an outcome to assume. If the passage repeats what stronger sources already say, or a bot cannot fetch it, your count of “gaps closed” can rise while citations stay flat.
Write the question at the top of the sheet before anyone drafts a page: When we fix a cited-source gap, does citation rate move more than it does for comparable gaps we leave alone? Everything below exists to answer that question.
Setup: freeze the prompts before you write
Build 30 to 40 buyer prompts from real customer questions, not a tool’s autocomplete. Pull from sales calls, support tickets, and the comparison pages your prospects read. Keep the exact wording. Sort prompts into discovery, use-case, comparison, problem, proof, and branded buckets, following the panel structure described in this guide to measuring AI brand visibility.
Make one row per prompt and engine. Record the wording, bucket, engine, date, each answer, whether your brand was mentioned, whether your domain was cited, and the cited URL. Keep the numerator and denominator visible. A prompt cited once in five runs is 1/5, not a mysterious visibility score.
Run each prompt five times per engine across ChatGPT, Perplexity, Gemini, and AI Overviews. Repetition matters: one practitioner baseline found that only 30% of brands persisted across back-to-back answers and 20% across five runs. For the fixed-panel approach behind this setup, see groas’s How to Measure AI Search Visibility Without Trusting One Screenshot. One screenshot is a draw. Five runs give you a rate you can compare later.
Establish the noise floor before making a change
Run the full baseline twice, one week apart, without changing the site between runs. Keep the prompts, engines, and run counts the same. For each prompt and engine, record citation rate as 0/5 through 5/5 and note how far it moves between baselines. That movement is your first read on the noise floor. Do not mistake a later 1/5-to-2/5 change for a breakthrough if it happened before you edited anything.
Rank-style tracking will not rescue a weak baseline. In the SparkToro and Gumshoe test of 2,961 repeated runs, fewer than 1 in 100 prompt pairs returned the same brand list, and fewer than 1 in 1,000 returned the same list in the same order. Count how often a citation appears; do not build your verdict around its position in one answer. The discipline matches AIVO’s advice to establish natural movement per metric before crediting a change.
That baseline is a warning threshold, not a magic number to apply to every result. You will compare treatment and control groups at the end. For now, it tells you how easily a prompt can appear to improve while nothing on your site changes.
Identify the passage your page does not answer
For prompts where you score 0/5 or 1/5, write down the cited competitor domain and exact URL. Open the page. Then write one sentence naming what its cited passage answers that yours does not. Name the passage, not the broad topic. “Their page gives a price threshold for a 20-person team; mine says pricing varies.” “Their page names the two tools it integrates with; mine says it integrates with your stack.”
If you cannot write that sentence, do not put the prompt into treatment yet. You have a missing citation, but you have not identified a fix. Keep looking or leave it out of this test. Publishing against a vague topic gap makes a later failure impossible to interpret: was the passage poor, or was it never the missing answer?
Assignment: hold half the clean gaps back
Choose 15 to 20 prompts with the clearest citation gaps from your frozen panel. Record the baseline for each, then randomly assign roughly half to treatment and half to control. A coin flip will do. Check the resulting groups against your prompt buckets and engines so you can see if one side happens to contain most of a particular kind of question. Keep the assignment and any imbalance in the sheet; do not quietly swap prompts after seeing an attractive target.
Publish fixes for treatment only. Touch nothing for control. If two prompts depend on the same page and fixing one would also fix the other, they are not independent holdouts. Keep them on the same side or choose different prompts before you publish. Otherwise the control is only a control in your spreadsheet.
Pinterest Engineering describes SEO experiments that separate enabled pages from controls and compare their change over time. That is the useful principle here; do not borrow its timeline as a promise for AI citations. I used to tell clients to fix everything at once. I was wrong. Fix everything and you learn nothing. Fix half and you learn what to buy more of.
Intervention: publish one answer, then check the fetch
For each treatment prompt, publish one answer-first passage. Give the missing answer in the first two sentences, then add the context a buyer needs to act on it. One prompt, one passage. Do not rewrite the whole site while you are trying to measure a small intervention.
A short page or a tight section on an existing page is enough for this protocol. Include the relevant numbers, thresholds, or names the cited passage supplies and yours lacks. Do not copy the competitor’s wording; answer the question directly. The structure follows groas’s passage-first guide to getting cited in AI Overviews and ChatGPT: direct answer, then proof, no preamble. Record the URL and publication date beside its treatment prompt.
Now check delivery. Open the live URL’s source and look for the answer in the returned HTML rather than trusting the CMS preview. Check server logs for successful requests from relevant bot user-agents. If you need a walkthrough, use groas’s Stop Rewriting for AI Citations. Check Whether Bots Can Read Your Page. The crawler visibility check is useful here too. A browser rendering a passage does not establish that a crawler received it.
A successful fetch is a prerequisite, not a citation. Log what you can verify. If the passage is absent from the response you inspect, or you cannot confirm the relevant bot fetched the page, flag that treatment. Do not count a CMS preview as a completed intervention and then blame the answer engine for ignoring it.
Measurement schedule: repeat the same runs
Use the first baseline as week 0 and the unchanged second baseline as week 1. Publish the treatment passages after that second baseline. Re-run the frozen prompts at weeks 2, 4, and 6, counted from week 0. Keep the wording, engines, five runs per engine, and recording method unchanged.
For each prompt, capture:
- Citation rate: the number of runs citing your domain out of five, recorded separately for each engine.
- Mention rate: the number naming your brand, whether or not they link to it.
- Winning URL: the page each answer cites, so you can see whether your new passage appears or a different page wins.
Touch nothing in control. If sales needs an edit to a control page mid-test, make the edit if you must, but log it and remove the affected prompt from the clean comparison. Do the same when a treatment passage changes materially after publication. A contaminated row does not become valid because deleting it makes the result less exciting.
Score engines separately before you combine anything. This ChatGPT-versus-Perplexity comparison reports different citation patterns across the two, including limited overlap in cited domains. If a passage gains citations in Perplexity and not ChatGPT, record that split. One blended visibility score would hide it. The useful tool here is the one that preserves prompts, engines, runs, and URLs instead of compressing them into a single grade.
Readout: compare gains, not screenshots
At week 6, calculate each prompt’s change in citation rate from baseline. Average those changes for the treatment prompts, then do the same for control. Subtract the control change from the treatment change. If treatment rose and control rose just as much, the treatment has not earned the credit. Compare the difference with the movement you observed during the unchanged baseline week, and inspect the individual prompt and engine rows before declaring a result.
This is a small, practical holdout, not a machine for proving every page caused every citation. Its job is to stop you treating ordinary answer variation as a paid-for win. Tape this line to the monitor: a lift that does not beat the holdout by more than the noise is not a lift.

If both groups rise together, do not write the case study. A model update, a competitor going offline, or a news cycle could have reshuffled sources. The holdout caught the false positive that a simple before-and-after would have sold you. Re-baseline before you credit the new pages.
If nothing moves, investigate in order:
- Fetch: Did a relevant bot request the live page, and was the answer present in the response you checked? No confirmed fetch means you cannot yet judge the passage.
- Eligibility versus selection: Allowing a crawler in
robots.txtor getting a page indexed does not mean an answer will cite it. In the Ahrefs matched-control test described here, pages with added JSON-LD did not show a clear citation gain against controls. A technical step can make a page available without making it the winning source. - Prompt and passage fit: Does the published answer address the question as buyers actually ask it? If not, you may have closed a gap on the page without closing the gap in the answer.
Do not use that checklist to declare a winner you cannot see. Use it to decide whether the next action is a delivery fix, a better passage, or a better prompt panel.
Decision: buy more fixes or stop buying gap lists
A gap list is a hypothesis, not a diagnosis. If treatment beats control beyond the movement you saw at baseline, you have a reason to repeat the kind of fix that worked and hold the vendor to the same test next time. Keep the winning prompts and engine-specific results; “AI visibility improved” is less useful than knowing which passages earned citations where.
If both groups move together, do not pay for credit the pages have not earned. If neither moves, check fetch and passage fit before commissioning more content. That is the use I have for gap platforms: prompt discovery and tracking, provided they let you keep the runs and cited URLs. The work is publishing a fetchable answer to a real buyer question and measuring whether it wins.
That is why I am blunt with owners about earned search. groas puts the passage work, fetch check, and re-measurement in one operation instead of handing you a spreadsheet of gaps and wishing you luck. The holdout still gets the final say.
Skip this test if you have fewer than 15 clean gaps, your site blocks the crawlers you need to check, or your category turns over weekly with news. Otherwise, spend the six weeks. Either the fixed half pulls ahead and you have a way to decide what to publish next, or it does not and you have a reason to stop paying to close gaps on faith. Both answers beat another dashboard.
Frequently asked questions
Is a screenshot of missing AI citations enough evidence to justify buying new pages to close gaps?
No. A grid of prompts where competitors are cited and your brand is not is only a hypothesis that closing the gaps will change the answers. To prove it, split the prompts in half, fix one half, hold the other back, and compare citation rates over six weeks. If the fixed half does not beat the holdout beyond normal movement, you bought pages, not visibility.
What is the difference between a visibility gap and a citation gap?
A Visibility Gap is a prompt where a competitor is mentioned and you are not, while a Citation Gap means a competitor domain is cited and yours is not. For the holdout test, use citation gaps, because a mention may point to where to investigate but cannot show whether a new passage earns a citation.
How should I build the prompt panel for an AI citation holdout test?
Build 30 to 40 buyer prompts from real customer questions in sales calls, support tickets, and comparison pages, keeping the exact wording, and sort them into discovery, use-case, comparison, problem, proof, and branded buckets. Run each prompt five times per engine across ChatGPT, Perplexity, Gemini, and AI Overviews, recording one row per prompt and engine so citation rates like 1/5 stay visible.
Why should I run the baseline twice before fixing anything?
Running the full baseline twice, one week apart and with nothing changed on the site, shows how much citation rates move on their own. That movement is your noise floor, and it keeps you from mistaking a 1/5-to-2/5 drift for a breakthrough caused by your edits. It is a warning threshold, not a fixed number applied to every result.
When should a prompt with a missing citation be left out of the treatment group?
Leave it out if you cannot write one sentence naming what the cited competitor passage answers that your page does not, at the level of a specific passage rather than a broad topic. Publishing against a vague topic gap makes a later failure impossible to interpret, because you cannot tell whether the passage was poor or never the missing answer.
How do I split prompts between treatment and control in the holdout test?
Choose 15 to 20 prompts with the clearest citation gaps, record each baseline, and randomly assign roughly half to treatment and half to control. Publish fixes only for treatment and touch nothing for control. If two prompts depend on the same page so fixing one would fix the other, keep them on the same side or pick different prompts, and do not quietly swap prompts after seeing an attractive target.
Do I need to verify that bots can fetch my new passage before measuring citations?
Yes. A successful fetch is a prerequisite, not a citation. Check the live URL's returned HTML rather than trusting the CMS preview, and check server logs for requests from relevant bot user-agents. If the passage is absent from the response or you cannot confirm the relevant bot fetched the page, flag that treatment rather than counting it as a completed intervention.
How do I read the results of the six-week holdout test?
At week 6, calculate each prompt's change in citation rate from baseline, average the changes for treatment and control separately, and subtract the control change from the treatment change. Compare that difference with the movement you saw during the unchanged baseline week. If treatment rose no more than control, the fix has not earned the credit, and you should re-baseline before crediting the new pages.




