A page gets rewritten on Monday. Perplexity cites it on Tuesday. By Friday, someone has a case study about the rewrite; by the following Tuesday, the citation is gone.
I would not accept that as evidence in an ad-copy test, and I would not accept it here. Most advice on getting cited by AI rests on the same unmonitored before-and-after story: change a page, spot a footnote, give the change credit. This protocol uses a holdout instead. Edit half a matched set of pages, leave the other half alone, and watch crawler fetches before deciding what the citations mean.
Question: did the edit earn a citation?
If you want to know how to get your business cited as a source in ChatGPT and Perplexity answers, or how to get your website cited in Google’s AI Overviews, isolate the edit from the engine’s ordinary churn. The test question is narrower than “Did our AI visibility improve?” It is: did one specified page change improve citation share relative to comparable pages you did not change?
The distinction matters. A GetMentions study of more than 530,000 citations found that 69% of an average answer’s cited sources changed day to day. In an Authoritas study of 11,203 keywords, Google AI Overviews had a volatility score of 0.68 to 0.73, versus 0.49 to 0.55 for traditional organic positions. Profound Strategy reported month-over-month citation drift of 59.3% in Google AI Overviews and 54.1% in ChatGPT. Tuesday’s footnote needs a better explanation than Tuesday’s FAQ accordion.
I would borrow the matched-pair logic used in search experiments, such as the framework in SearchPilot’s SEO A/B testing guide. This is not a clean user-level split test: AI engines can change their retrieval and answers while you run it. But a set of untouched pages gives you a way to see whether edited pages moved differently from similar pages exposed to the same period of churn.
Setup: choose prompts, pair pages, then leave them alone
Record 10–20 buyer prompts
Start with the searches that could plausibly lead to pipeline. Skip broad definitions such as “what is inventory management” and choose specific questions your pages can answer:
- Comparisons: “Tool A vs Tool B for mid-market field service dispatch”
- Scope and pricing: “average onboarding cost for warehouse management systems”
- Implementation: “how to integrate custom inventory software with netsuite”
Put the exact prompt strings in a spreadsheet. Assign each prompt to the page or matched pair it is meant to test. Keep wording and punctuation fixed between scheduled runs. Record which engine you used and how you accessed it; do not mix a logged-in colleague’s casual searches into the test sessions. Otherwise, a changed prompt or browsing context becomes another explanation for a changed answer.
Match pages before touching the HTML
Select 10 to 20 candidate URLs that address your buyer prompts. Pair pages on topic, page age, and baseline 90-day organic impressions from Google Search Console. A fleet-tracking pricing breakdown might pair with an onboarding-cost breakdown for the same industry. They will not be identical, but they should be close enough that an engine-wide change affecting the topic has a chance to show up in both.
Within each pair, assign one URL to Group A, the edited variant, and the other to Group B, the holdout. Keep Group B untouched: no copy, schema, or internal-link edits during the run. Document any sitewide changes that affect both groups. If the pages differ sharply in existing visibility, find a better pair rather than expecting arithmetic to rescue the comparison later.

Establish a baseline for each engine
Log at least 14 days before the intervention. Run each prompt at least five separate times per engine across that window in ChatGPT Search, Perplexity, and Google Search, noting whether an AI Overview appears. For every run, record:
- Whether your domain appears in a citation, footnote, or card.
- The exact URL cited, especially whether it belongs to Group A or Group B.
- The other domains cited.
- The prompt, engine, date, and session conditions.
The exact URL matters more than a domain-level victory lap. If an engine cites your homepage instead of the page you edited, the edit has not earned credit. As I argued in this breakdown of measuring AI search visibility, one screenshot cannot establish a trend. Use the baseline runs to learn how often each page appears before you change it.
Check whether relevant bots reach the pages
Before editing, inspect server access logs or CDN edge data for requests to each test URL. Keep background crawling separate from live-answer activity where the user agent lets you do so. OpenAI distinguishes ChatGPT-User from OAI-SearchBot and GPTBot; Perplexity documents Perplexity-User. You may also see traffic from agents outside the three products in this citation test, including Anthropic’s Claude-User or Google-Agent. Log them separately rather than treating every AI-labelled request as a fetch for one of your test answers.
A bot fetch is a clue, not a citation. A request does not tell you which prompt caused it, and an answer may draw on material fetched earlier. Still, logs can expose a basic obstacle: if relevant requests never reach either group, check whether bots can read the pages before spending another sprint on headings. Check response status and timing too. A page that fails or responds too slowly is a different problem from a page an engine reads and declines to cite.
Intervention: change one thing on Group A
Choose one variable and apply the same kind of edit across the Group A pages. Leave Group B alone until the test ends. Options include:
- Answer-first copy: Rewrite only the opening 150 words of each target section so it answers the assigned buyer prompt directly. Leave titles, schema, and page structure untouched.
- Entity clarity: Replace ambiguous brand or feature references with explicit entity-attribute-value statements in the body. Do not also change headings.
- Structured data: Add the chosen JSON-LD schema while leaving visible HTML unchanged.
Write down what changed, which URLs received it, and when it went live. Do not submit a new sitemap, refresh metadata, add an FAQ, and rewrite the copy in the same sprint unless your question is whether that entire bundle works. It will not tell you which part mattered.
That restraint is useful because popular tactics do not always survive isolation. In an Ahrefs investigation comparing 1,885 pages that added JSON-LD with 4,000 controls over 30 days, schema showed no meaningful citation uplift in ChatGPT or Google AI Mode; Google AI Overviews showed a 4.6% decline consistent with pre-intervention baselines. The same investigation found that 97% of llms.txt files across 137,000 sites were never fetched by AI bots. If you change schema, prose, and llms.txt together, a later citation will not tell you which change deserves credit. It may tell you nothing about any of them.
Measurement: fetches first, citations next
Watch requests without mistaking them for results
Check per-URL bot requests daily after publication. Compare Group A with Group B and with each group’s own baseline. Separate identifiable live-fetch user agents from indexing crawls, and record successful responses as well as failures.
A sustained difference in fetches is an early signal worth investigating, not proof that the edit won. More requests to Group A could mean its pages entered more retrieval paths. It could also reflect crawling unrelated to your scheduled prompts. Conversely, unchanged live-fetch counts do not prove the edit was never evaluated: an engine may use indexed or cached material. Use logs to locate a possible mechanism, then look for the citation outcome.

Calculate citation share for the page you tested
Across the post-intervention window, schedule ten distinct sessions per prompt on each engine. Record the actual cited URLs in ChatGPT Search and Perplexity answers and, when Google shows an AI Overview, in its citation card. Do not substitute an organic ranking report for that card. The Authoritas study found that 40% of the time, AI Overviews cited a webpage outside Google’s organic top 10.
Citation share is the percentage of relevant runs that cite the target page. If it appears in seven of ten runs, its share for that prompt is 70%. Track domain citations separately so another page on your site does not inflate the result for the edited URL. Keep engine results separate as well: a change that appears to help in Perplexity has not automatically helped in ChatGPT or AI Overviews.

Subtract the holdout’s movement
Now compare changes, not just final percentages. Suppose both groups average 20% citation share during the baseline. In the evaluation window, Group B rises to 26% without an edit, while Group A rises to 34%. Group A’s raw increase is 14 percentage points; its increase relative to the holdout is eight points. That eight-point difference is the signal to examine, not a guaranteed causal effect.
If Group A rises five points and Group B rises six, the rewrite has not beaten the background movement. Inspect individual pairs as well as the group average. One unusually strong page should not carry nine weak ones into a triumphant slide.
Readout: win, no effect, or another test
Set a decision rule before you inspect the post-edit answers. For this protocol, I would look for a sustained change in relevant fetches, then require Group A’s citation-share gain to exceed Group B’s by at least 15 percentage points across a two-week evaluation window. That is a deliberately demanding working threshold, not a universal law or a statistical guarantee. Report the number of runs and pairs beside any result.
- A promising win: Group A gains relevant fetch activity and beats Group B’s citation-share change by the predeclared margin. Check that the lift appears across multiple pairs and is not confined to one engine before rolling out the edit.
- No observed effect: Fetch activity stays broadly similar and citation share moves within the variation seen during the baseline, with no meaningful advantage over Group B. Keep the current structure; this test has not justified a wider rollout.
- Inconclusive: Group A gets more fetches but no citation-share lift, or citations rise without a comparable holdout advantage. Retrieval may have changed while source selection did not, but the logs cannot prove what the answer-generation step preferred. Inspect the runs and repeat rather than naming a winner.
Do not run this on two URLs and three prompt checks. With the answer churn documented above, one changed citation would dominate that tiny sample. Use at least five matched pairs and aim for 100 recorded prompt sessions across the tested engines over the 14-day evaluation window. These are operating minimums for a readable comparison, not a promise that the result will be conclusive. If you cannot sustain the logging schedule, do not turn a handful of screenshots into a causal claim.
When to skip the protocol
Skip it if you cannot assemble reasonably matched pages, keep a group untouched, or access logs or edge metrics that identify requests to the tested URLs. Without those controls, you can still observe citations, but you cannot run this test as written.
I would also hold off if the site has little search visibility and the target pages rarely appear for related topics. First check that the pages are discoverable and that you can establish a usable citation baseline. A holdout cannot extract a clean signal from pages the engines never encounter. That is not an argument for rewriting everything until a bot shows up. It is an argument for fixing the measurement setup before calling the rewrite a success or failure.
What the result changes
If the edited pages beat the holdouts by the rule you set in advance, extend that one change to the rest of the topic cluster and keep measuring. If they do not, leave the untouched pages alone and test a different variable. If fetches change but citations do not, investigate what the engines can retrieve and what they choose to cite before spending more time on copy.
Then ask the question a citation screenshot cannot answer: did that visibility contribute to qualified pipeline or revenue? I spent years watching teams celebrate impression share while cost per acquisition quietly worsened. A footnote can become the organic-search version of the same mistake.
That is why groas connects paid and organic search execution to attributable outcomes rather than treating a citation count as the finish line. Its autonomous execution works within client guardrails, with a named strategist accountable for direction. Whether you use that model or run the protocol yourself, the next decision should follow the result: roll out a change that beats the holdout, reject one that does not, and stop giving a Tuesday footnote credit for work it may never have seen.

