---
title: "Did Your Rewrite Earn That AI Citation? Run a Holdout Test"
description: "Test whether a page edit earns citations in ChatGPT, Perplexity, or Google AI Overviews. Compare matched pages, crawler fetches, and citation share against a holdout."
url: "https://groas.com/post/did-that-page-change-actually-get-you-ci"
image: "https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/b9245ddc-b454-4fc7-b2cb-14455750d407.png"
published: "2026-10-08T05:49:00.195Z"
modified: "2026-10-08T05:49:00.261Z"
---

October 8, 2026 · 10 min read

# Did Your Rewrite Earn That AI Citation? Run a Holdout Test

[Alexander PerelmanHead Of Product @ groas](https://groas.com/author/alexander-perelman)

![Cartoon gardener proudly credits his watering can for a lush bed, while rain falls on both beds and the untouched one beside it grew just as much.](https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/b9245ddc-b454-4fc7-b2cb-14455750d407.png)

In this article

1. [Question: did the edit earn a citation?](#question-did-the-edit-earn-a-citation)
2. [Setup: choose prompts, pair pages, then leave them alone](#setup-choose-prompts-pair-pages-then-leave-them-alone)
3. [Intervention: change one thing on Group A](#intervention-change-one-thing-on-group-a)
4. [Measurement: fetches first, citations next](#measurement-fetches-first-citations-next)
5. [Readout: win, no effect, or another test](#readout-win-no-effect-or-another-test)
6. [When to skip the protocol](#when-to-skip-the-protocol)
7. [What the result changes](#what-the-result-changes)

A page gets rewritten on Monday. Perplexity cites it on Tuesday. By Friday, someone has a case study about the rewrite; by the following Tuesday, the citation is gone.

I would not accept that as evidence in an ad-copy test, and I would not accept it here. Most advice on getting cited by AI rests on the same unmonitored before-and-after story: change a page, spot a footnote, give the change credit. This protocol uses a holdout instead. Edit half a matched set of pages, leave the other half alone, and watch crawler fetches before deciding what the citations mean.

## Question: did the edit earn a citation?

If you want to know **how to get your business cited as a source in ChatGPT and Perplexity answers**, or **how to get your website cited in Google’s AI Overviews**, isolate the edit from the engine’s ordinary churn. The test question is narrower than “Did our AI visibility improve?” It is: did one specified page change improve citation share _relative to comparable pages you did not change_?

The distinction matters. A [GetMentions study of more than 530,000 citations](https://getmentions.ai/blog/ai-citation-volatility-study) found that 69% of an average answer’s cited sources changed day to day. In an [Authoritas study of 11,203 keywords](https://searchengineland.com/google-ai-overviews-more-volatile-than-organic-rankings-451877), Google AI Overviews had a volatility score of 0.68 to 0.73, versus 0.49 to 0.55 for traditional organic positions. [Profound Strategy reported month-over-month citation drift](https://profoundstrategy.com/research/ai-search-volatility/) of 59.3% in Google AI Overviews and 54.1% in ChatGPT. Tuesday’s footnote needs a better explanation than Tuesday’s FAQ accordion.

I would borrow the matched-pair logic used in search experiments, such as the framework in [SearchPilot’s SEO A/B testing guide](https://www.searchpilot.com/resources/blog/what-is-seo-ab-testing-guide/). This is not a clean user-level split test: AI engines can change their retrieval and answers while you run it. But a set of untouched pages gives you a way to see whether edited pages moved differently from similar pages exposed to the same period of churn.

## Setup: choose prompts, pair pages, then leave them alone

### Record 10–20 buyer prompts

Start with the searches that could plausibly lead to pipeline. Skip broad definitions such as “what is inventory management” and choose specific questions your pages can answer:

- **Comparisons:** “Tool A vs Tool B for mid-market field service dispatch”
- **Scope and pricing:** “average onboarding cost for warehouse management systems”
- **Implementation:** “how to integrate custom inventory software with netsuite”

Put the exact prompt strings in a spreadsheet. Assign each prompt to the page or matched pair it is meant to test. Keep wording and punctuation fixed between scheduled runs. Record which engine you used and how you accessed it; do not mix a logged-in colleague’s casual searches into the test sessions. Otherwise, a changed prompt or browsing context becomes another explanation for a changed answer.

### Match pages before touching the HTML

Select 10 to 20 candidate URLs that address your buyer prompts. Pair pages on **topic, page age, and baseline 90-day organic impressions** from Google Search Console. A fleet-tracking pricing breakdown might pair with an onboarding-cost breakdown for the same industry. They will not be identical, but they should be close enough that an engine-wide change affecting the topic has a chance to show up in both.

Within each pair, assign one URL to **Group A, the edited variant**, and the other to **Group B, the holdout**. Keep Group B untouched: no copy, schema, or internal-link edits during the run. Document any sitewide changes that affect both groups. If the pages differ sharply in existing visibility, find a better pair rather than expecting arithmetic to rescue the comparison later.

![Illustration of two filing cards balanced on a scale beneath a magnifying lens.](https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/a8602a21-b670-4815-8163-13e2b8c5877d.png)

### Establish a baseline for each engine

Log at least **14 days before the intervention**. Run each prompt at least five separate times per engine across that window in ChatGPT Search, Perplexity, and Google Search, noting whether an AI Overview appears. For every run, record:

1. Whether your domain appears in a citation, footnote, or card.
2. The exact URL cited, especially whether it belongs to Group A or Group B.
3. The other domains cited.
4. The prompt, engine, date, and session conditions.

The exact URL matters more than a domain-level victory lap. If an engine cites your homepage instead of the page you edited, the edit has not earned credit. As I argued in [this breakdown of measuring AI search visibility](https://groas.com/post/how-to-measure-your-brand-s-visibility-i-2), one screenshot cannot establish a trend. Use the baseline runs to learn how often each page appears _before_ you change it.

### Check whether relevant bots reach the pages

Before editing, inspect server access logs or CDN edge data for requests to each test URL. Keep background crawling separate from live-answer activity where the user agent lets you do so. OpenAI distinguishes [`ChatGPT-User` from `OAI-SearchBot` and `GPTBot`](https://developers.openai.com/docs/bots); Perplexity documents [`Perplexity-User`](https://docs.perplexity.ai/guides/bots). You may also see traffic from agents outside the three products in this citation test, including [Anthropic’s `Claude-User`](https://support.anthropic.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) or [`Google-Agent`](https://searchengineland.com/google-agent-user-agent-server-logs-453120). Log them separately rather than treating every AI-labelled request as a fetch for one of your test answers.

**A bot fetch is a clue, not a citation.** A request does not tell you which prompt caused it, and an answer may draw on material fetched earlier. Still, logs can expose a basic obstacle: if relevant requests never reach either group, [check whether bots can read the pages](https://groas.com/post/stop-rewriting-your-content-for-ai-your) before spending another sprint on headings. Check response status and timing too. A page that fails or responds too slowly is a different problem from a page an engine reads and declines to cite.

## Intervention: change one thing on Group A

Choose **one variable** and apply the same kind of edit across the Group A pages. Leave Group B alone until the test ends. Options include:

- **Answer-first copy:** Rewrite only the opening 150 words of each target section so it answers the assigned buyer prompt directly. Leave titles, schema, and page structure untouched.
- **Entity clarity:** Replace ambiguous brand or feature references with explicit entity-attribute-value statements in the body. Do not also change headings.
- **Structured data:** Add the chosen JSON-LD schema while leaving visible HTML unchanged.

Write down what changed, which URLs received it, and when it went live. Do not submit a new sitemap, refresh metadata, add an FAQ, and rewrite the copy in the same sprint unless your question is whether that entire bundle works. It will not tell you which part mattered.

That restraint is useful because popular tactics do not always survive isolation. In an [Ahrefs investigation comparing 1,885 pages that added JSON-LD with 4,000 controls over 30 days](https://ahrefs.com/blog/ai-search-myths/), schema showed no meaningful citation uplift in ChatGPT or Google AI Mode; Google AI Overviews showed a 4.6% decline consistent with pre-intervention baselines. The same investigation found that 97% of `llms.txt` files across 137,000 sites were never fetched by AI bots. If you change schema, prose, and `llms.txt` together, a later citation will not tell you which change deserves credit. It may tell you nothing about any of them.

## Measurement: fetches first, citations next

### Watch requests without mistaking them for results

Check per-URL bot requests daily after publication. Compare Group A with Group B and with each group’s own baseline. Separate identifiable live-fetch user agents from indexing crawls, and record successful responses as well as failures.

**A sustained difference in fetches is an early signal worth investigating, not proof that the edit won.** More requests to Group A could mean its pages entered more retrieval paths. It could also reflect crawling unrelated to your scheduled prompts. Conversely, unchanged live-fetch counts do not prove the edit was never evaluated: an engine may use indexed or cached material. Use logs to locate a possible mechanism, then look for the citation outcome.

![Cutaway diagram of server logs separating AI bot requests by response status.](https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/637e4533-fc54-4479-8edf-743294fe04fc.png)

### Calculate citation share for the page you tested

Across the post-intervention window, schedule **ten distinct sessions per prompt on each engine**. Record the actual cited URLs in ChatGPT Search and Perplexity answers and, when Google shows an AI Overview, in its citation card. Do not substitute an organic ranking report for that card. The [Authoritas study](https://searchengineland.com/google-ai-overviews-more-volatile-than-organic-rankings-451877) found that 40% of the time, AI Overviews cited a webpage outside Google’s organic top 10.

Citation share is the percentage of relevant runs that cite the target page. If it appears in seven of ten runs, its share for that prompt is 70%. Track domain citations separately so another page on your site does not inflate the result for the edited URL. Keep engine results separate as well: a change that appears to help in Perplexity has not automatically helped in ChatGPT or AI Overviews.

![Three-panel chart comparing raw citation lift, holdout lift, and baseline noise.](https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/9d4a6c10-11c7-4926-9305-be990704c705.png)

### Subtract the holdout’s movement

Now compare changes, not just final percentages. Suppose both groups average 20% citation share during the baseline. In the evaluation window, Group B rises to 26% without an edit, while Group A rises to 34%. Group A’s raw increase is 14 percentage points; its increase _relative to the holdout_ is eight points. That eight-point difference is the signal to examine, not a guaranteed causal effect.

If Group A rises five points and Group B rises six, the rewrite has not beaten the background movement. Inspect individual pairs as well as the group average. One unusually strong page should not carry nine weak ones into a triumphant slide.

## Readout: win, no effect, or another test

Set a decision rule **before** you inspect the post-edit answers. For this protocol, I would look for a sustained change in relevant fetches, then require Group A’s citation-share gain to exceed Group B’s by at least 15 percentage points across a two-week evaluation window. That is a deliberately demanding working threshold, not a universal law or a statistical guarantee. Report the number of runs and pairs beside any result.

- **A promising win:** Group A gains relevant fetch activity and beats Group B’s citation-share change by the predeclared margin. Check that the lift appears across multiple pairs and is not confined to one engine before rolling out the edit.
- **No observed effect:** Fetch activity stays broadly similar and citation share moves within the variation seen during the baseline, with no meaningful advantage over Group B. Keep the current structure; this test has not justified a wider rollout.
- **Inconclusive:** Group A gets more fetches but no citation-share lift, or citations rise without a comparable holdout advantage. Retrieval may have changed while source selection did not, but the logs cannot prove what the answer-generation step preferred. Inspect the runs and repeat rather than naming a winner.

Do not run this on two URLs and three prompt checks. With the answer churn documented above, one changed citation would dominate that tiny sample. Use **at least five matched pairs** and aim for **100 recorded prompt sessions** across the tested engines over the 14-day evaluation window. These are operating minimums for a readable comparison, not a promise that the result will be conclusive. If you cannot sustain the logging schedule, do not turn a handful of screenshots into a causal claim.

## When to skip the protocol

Skip it if you cannot assemble reasonably matched pages, keep a group untouched, or access logs or edge metrics that identify requests to the tested URLs. Without those controls, you can still observe citations, but you cannot run _this_ test as written.

I would also hold off if the site has little search visibility and the target pages rarely appear for related topics. First check that the pages are discoverable and that you can establish a usable citation baseline. A holdout cannot extract a clean signal from pages the engines never encounter. That is not an argument for rewriting everything until a bot shows up. It is an argument for fixing the measurement setup before calling the rewrite a success or failure.

## What the result changes

If the edited pages beat the holdouts by the rule you set in advance, extend that one change to the rest of the topic cluster and keep measuring. If they do not, leave the untouched pages alone and test a different variable. If fetches change but citations do not, investigate what the engines can retrieve and what they choose to cite before spending more time on copy.

Then ask the question a citation screenshot cannot answer: did that visibility contribute to qualified pipeline or revenue? I spent years watching teams celebrate impression share while cost per acquisition quietly worsened. A footnote can become the organic-search version of the same mistake.

That is why [groas](https://groas.com/) connects paid and organic search execution to attributable outcomes rather than treating a citation count as the finish line. Its autonomous execution works within client guardrails, with a named strategist accountable for direction. Whether you use that model or run the protocol yourself, the next decision should follow the result: roll out a change that beats the holdout, reject one that does not, and stop giving a Tuesday footnote credit for work it may never have seen.

## Frequently Asked Questions

### Why can't I just compare citations before and after a page rewrite to see if the rewrite worked?

Because AI answers change constantly on their own. A GetMentions study found 69% of an average answer's cited sources change day to day, and Google AI Overviews have a volatility score of 0.68 to 0.73 versus 0.49 to 0.55 for organic rankings. A holdout group of untouched pages lets you see whether edited pages moved differently from comparable pages exposed to the same background churn.

### How do I set up matched pages for an AI citation holdout test?

Select 10 to 20 candidate URLs and pair them on topic, page age, and baseline 90-day organic impressions from Google Search Console. Within each pair, assign one URL to Group A, the edited variant, and the other to Group B, the holdout. Keep Group B untouched, with no copy, schema, or internal-link edits, and document any sitewide changes affecting both groups.

### How long should I record a baseline before making the page edit?

Log at least 14 days before the intervention and run each prompt at least five separate times per engine in ChatGPT Search, Perplexity, and Google Search. For every run, record whether your domain is cited, the exact URL cited, the other domains cited, and the prompt, engine, date, and session conditions.

### Should I look at AI bot fetches in my server logs during the test?

Yes, but treat a bot fetch as a clue rather than proof. Check per-URL requests daily, separating live-fetch user agents such as OAI-SearchBot and Perplexity-User from indexing crawls, and compare Group A against Group B and each baseline. A sustained difference in fetches is an early signal worth investigating, but the citation outcome is the real result.

### How many changes should I make to the edited pages during the test?

Choose one variable and apply the same kind of edit across the Group A pages, such as answer-first opening copy, entity clarity, or a single JSON-LD schema addition. Changing copy, schema, and llms.txt together makes it impossible to know which change earned any citation. Note that an Ahrefs investigation found schema showed no meaningful citation uplift in ChatGPT or Google AI Mode.

### What is citation share and how do I calculate it?

Citation share is the percentage of relevant runs that cite the target page. If the page appears in seven of ten scheduled runs for a prompt, its share for that prompt is 70%. Track the exact URL cited rather than just the domain, and keep results separate per engine and per prompt.

### How much improvement does an edited page need to count as a win over the holdout?

Compare the change in citation share, not just final percentages. In this protocol, Group A's citation-share gain must exceed Group B's by at least 15 percentage points across a two-week evaluation window, alongside a sustained change in relevant fetches. That threshold is set before inspecting results and is a deliberately demanding working rule, not a statistical guarantee.

### How many page pairs and test sessions does this experiment need?

Use at least five matched pairs and aim for 100 recorded prompt sessions across the tested engines over the 14-day evaluation window. Running the test on two URLs and three prompt checks lets a single changed citation dominate the sample. If you cannot sustain the logging schedule, do not turn a handful of screenshots into a causal claim.

## Pay For Results, Not For Hours

Businesses buy the outcome, agencies resell it, and groas answers for it either way.

[See If You Qualify](https://groas.typeform.com/to/xC1bQNUT)

## Structured data

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://graph.groas.com/entity/groas","name":"groas","url":"https://groas.com/","logo":{"@type":"ImageObject","url":"https://groas.com/icon-512.png","width":512,"height":512},"sameAs":["https://www.linkedin.com/company/groas/"]},{"@type":"WebSite","@id":"https://groas.com/#website","name":"groas","url":"https://groas.com/","inLanguage":"en-US","publisher":{"@id":"https://graph.groas.com/entity/groas"}},{"@type":"WebPage","@id":"https://groas.com/post/did-that-page-change-actually-get-you-ci#webpage","url":"https://groas.com/post/did-that-page-change-actually-get-you-ci","name":"Did Your Rewrite Earn That AI Citation? Run a Holdout Test","description":"Test whether a page edit earns citations in ChatGPT, Perplexity, or Google AI Overviews. Compare matched pages, crawler fetches, and citation share against a holdout.","inLanguage":"en-US","isPartOf":{"@id":"https://groas.com/#website"},"about":{"@id":"https://graph.groas.com/entity/groas"},"primaryImageOfPage":{"@type":"ImageObject","url":"https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/b9245ddc-b454-4fc7-b2cb-14455750d407.png"},"breadcrumb":{"@id":"https://groas.com/post/did-that-page-change-actually-get-you-ci#breadcrumb"},"dateModified":"2026-10-08T05:49:00.261Z","mainEntity":{"@id":"https://groas.com/post/did-that-page-change-actually-get-you-ci#article"}},{"@type":"BreadcrumbList","@id":"https://groas.com/post/did-that-page-change-actually-get-you-ci#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://groas.com/"},{"@type":"ListItem","position":2,"name":"Blog","item":"https://groas.com/blog"},{"@type":"ListItem","position":3,"name":"Did Your Rewrite Earn That AI Citation? Run a Holdout Test","item":"https://groas.com/post/did-that-page-change-actually-get-you-ci"}]},{"@type":"BlogPosting","@id":"https://groas.com/post/did-that-page-change-actually-get-you-ci#article","url":"https://groas.com/post/did-that-page-change-actually-get-you-ci","mainEntityOfPage":{"@id":"https://groas.com/post/did-that-page-change-actually-get-you-ci#webpage"},"isPartOf":{"@id":"https://groas.com/blog#blog"},"headline":"Did Your Rewrite Earn That AI Citation? Run a Holdout Test","description":"Test whether a page edit earns citations in ChatGPT, Perplexity, or Google AI Overviews. Compare matched pages, crawler fetches, and citation share against a holdout.","datePublished":"2026-10-08T05:49:00.195Z","dateModified":"2026-10-08T05:49:00.261Z","image":{"@type":"ImageObject","url":"https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/b9245ddc-b454-4fc7-b2cb-14455750d407.png"},"author":{"@id":"https://groas.com/author/alexander-perelman#person"},"publisher":{"@id":"https://graph.groas.com/entity/groas"},"keywords":"citation, group, test, pages, page, holdout, change, edit","wordCount":2122,"inLanguage":"en-US"},{"@type":"Person","@id":"https://groas.com/author/alexander-perelman#person","name":"Alexander Perelman","jobTitle":"Head Of Product @ groas","description":"Ex Goldman Sachs and Ex Stanford Computer Science","image":"https://groas.com/media/blog/4f17cac3a81acc224cc1b2cfab8f94bbc295a3089d13e8421125771ef8d4064f.jpg","worksFor":{"@id":"https://graph.groas.com/entity/groas"},"sameAs":["https://www.linkedin.com/in/alexander-433793253/"]},{"@type":"FAQPage","@id":"https://groas.com/post/did-that-page-change-actually-get-you-ci#faq","mainEntity":[{"@type":"Question","name":"Why can't I just compare citations before and after a page rewrite to see if the rewrite worked?","acceptedAnswer":{"@type":"Answer","text":"Because AI answers change constantly on their own. A GetMentions study found 69% of an average answer's cited sources change day to day, and Google AI Overviews have a volatility score of 0.68 to 0.73 versus 0.49 to 0.55 for organic rankings. A holdout group of untouched pages lets you see whether edited pages moved differently from comparable pages exposed to the same background churn."}},{"@type":"Question","name":"How do I set up matched pages for an AI citation holdout test?","acceptedAnswer":{"@type":"Answer","text":"Select 10 to 20 candidate URLs and pair them on topic, page age, and baseline 90-day organic impressions from Google Search Console. Within each pair, assign one URL to Group A, the edited variant, and the other to Group B, the holdout. Keep Group B untouched, with no copy, schema, or internal-link edits, and document any sitewide changes affecting both groups."}},{"@type":"Question","name":"How long should I record a baseline before making the page edit?","acceptedAnswer":{"@type":"Answer","text":"Log at least 14 days before the intervention and run each prompt at least five separate times per engine in ChatGPT Search, Perplexity, and Google Search. For every run, record whether your domain is cited, the exact URL cited, the other domains cited, and the prompt, engine, date, and session conditions."}},{"@type":"Question","name":"Should I look at AI bot fetches in my server logs during the test?","acceptedAnswer":{"@type":"Answer","text":"Yes, but treat a bot fetch as a clue rather than proof. Check per-URL requests daily, separating live-fetch user agents such as OAI-SearchBot and Perplexity-User from indexing crawls, and compare Group A against Group B and each baseline. A sustained difference in fetches is an early signal worth investigating, but the citation outcome is the real result."}},{"@type":"Question","name":"How many changes should I make to the edited pages during the test?","acceptedAnswer":{"@type":"Answer","text":"Choose one variable and apply the same kind of edit across the Group A pages, such as answer-first opening copy, entity clarity, or a single JSON-LD schema addition. Changing copy, schema, and llms.txt together makes it impossible to know which change earned any citation. Note that an Ahrefs investigation found schema showed no meaningful citation uplift in ChatGPT or Google AI Mode."}},{"@type":"Question","name":"What is citation share and how do I calculate it?","acceptedAnswer":{"@type":"Answer","text":"Citation share is the percentage of relevant runs that cite the target page. If the page appears in seven of ten scheduled runs for a prompt, its share for that prompt is 70%. Track the exact URL cited rather than just the domain, and keep results separate per engine and per prompt."}},{"@type":"Question","name":"How much improvement does an edited page need to count as a win over the holdout?","acceptedAnswer":{"@type":"Answer","text":"Compare the change in citation share, not just final percentages. In this protocol, Group A's citation-share gain must exceed Group B's by at least 15 percentage points across a two-week evaluation window, alongside a sustained change in relevant fetches. That threshold is set before inspecting results and is a deliberately demanding working rule, not a statistical guarantee."}},{"@type":"Question","name":"How many page pairs and test sessions does this experiment need?","acceptedAnswer":{"@type":"Answer","text":"Use at least five matched pairs and aim for 100 recorded prompt sessions across the tested engines over the 14-day evaluation window. Running the test on two URLs and three prompt checks lets a single changed citation dominate the sample. If you cannot sustain the logging schedule, do not turn a handful of screenshots into a causal claim."}}]}]}
```
