---
title: "Before You Rewrite for AI Citations, Run This 30-Day Server Log Test"
description: "Use a 30-day server log protocol to see whether ChatGPT and Perplexity fetch your pages, what they receive, and which fix to test before rewriting copy."
url: "https://groas.com/post/are-ai-assistants-even-reading-your-page"
image: "https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/5dd082f1-dd0f-4f61-9b46-8931ae33f61e.png"
published: "2026-10-11T05:38:28.522Z"
modified: "2026-10-11T05:38:28.597Z"
---

October 11, 2026 · 10 min read

# Before You Rewrite for AI Citations, Run This 30-Day Server Log Test

[Alexander PerelmanHead Of Product @ groas](https://groas.com/author/alexander-perelman)[LinkedIn](https://www.linkedin.com/in/alexander-433793253/)

![A shopkeeper on a ladder repaints a glossy green storefront while its front door is bricked shut, so no one can get in to see the new paint.](https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/5dd082f1-dd0f-4f61-9b46-8931ae33f61e.png)

In this article

1. [The question: Are assistants fetching your candidate pages?](#the-question-are-assistants-fetching-your-candidate-pages)
2. [Day 0: Collect the logs that can see a bot](#day-0-collect-the-logs-that-can-see-a-bot)
3. [Days 1–30: Collect three streams without changing the site](#days-130-collect-three-streams-without-changing-the-site)
4. [Day 30: Sort the results by failure mode](#day-30-sort-the-results-by-failure-mode)
5. [Follow-up: Change one lever per cohort](#follow-up-change-one-lever-per-cohort)
6. [Limits: A log is evidence, not an assistant transcript](#limits-a-log-is-evidence-not-an-assistant-transcript)
7. [One-page diagnostic scorecard](#one-page-diagnostic-scorecard)

Before you rewrite a page for ChatGPT or Perplexity citations, check whether either assistant has fetched it. If the request never reaches the page, a sharper headline will not help. That is repainting a storefront with the front door bricked shut.

We took a first look at the source-selection problem in [_Getting Cited by ChatGPT and Perplexity Is a Retrieval Auction_](https://groas.com/post/how-chatgpt-and-perplexity-actually-pick). This is the second look because knowing how assistants choose sources only helps if you can see which step fails on your own site. Your server and edge logs can show you the requests they receive. Run the 30-day test before you pay someone to rewrite the copy.

## The question: Are assistants fetching your candidate pages?

This is a prospective audit, not a promise that a log entry will explain every citation. The operational question is: **When you test buyer prompts, do verified live-retrieval bots request your candidate URLs, and what does your site return?**

My expectation is that many commercial sites will find uneven traffic: little or no live fetching of core money pages, with more activity on a few narrow guides or technical pages. That is a hypothesis to test, not a result to assume. A page has to be available as a candidate and useful to the assistant at the point of retrieval. Those are different problems, and the fixes should be different too.

## Day 0: Collect the logs that can see a bot

Forget Google Analytics or Plausible for this job. Client-side analytics depend on scripts that a retrieval agent may never run; [this field guide to AI crawlers in server logs](https://usegeon.com/blog/en/reading-ai-crawlers-in-your-server-logs-a-practical-field-guide) explains the distinction. Use raw NGINX or Apache access logs, or request logs from an edge provider such as Cloudflare, Fastly, or AWS CloudFront. Retain 30 days with timestamps, client IPs, full request URIs, HTTP status codes, and `User-Agent` strings. Capture response timing and edge actions where available.

**Keep edge and origin records separate.** A request blocked at the edge cannot appear in an origin log. If you inspect only NGINX while your WAF rejects the bot upstream, you may call it a zero-fetch page when the bot did try to visit.

![Diagram separating training crawlers, search indexers, and live user-action fetchers, with an IP verification step](https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/bf2f82f0-c464-48e3-a2aa-e7c9faa24157.png)

Next, split the traffic by job. [Cloudflare Radar’s AI crawler telemetry](https://blog.cloudflare.com/ai-crawler-traffic-by-purpose-and-industry/) reports that training crawlers account for roughly 80% of AI bot activity, while user-action crawlers account for less than 5%. Do not mix those populations in a citation test:

- **Training crawlers:** `GPTBot` and `ClaudeBot` crawl asynchronously. Their visits do not show that someone asked an assistant about your product at that moment.
- **Search indexers:** [OpenAI documents `OAI-SearchBot`](https://developers.openai.com/api/docs/bots), and [Perplexity documents `PerplexityBot`](https://docs.perplexity.ai/docs/resources/perplexity-crawlers). Track them separately from live user-action requests.
- **User-action fetchers:** `ChatGPT-User` and `Perplexity-User` are the request classes to inspect when you test live lookups. A fetch shows a request for the URL, not that the assistant read it successfully or cited it. Perplexity says `Perplexity-User` generally ignores robots.txt disallow directives because it acts on a user request.

A `User-Agent` is a label anyone can print. [HUMAN Security’s analysis of crawler spoofing](https://www.humansecurity.com/learn/blog/ai-crawler-spoofing-chatgpt-mistral-perplexity/) found unauthorized requests among traffic claiming to be AI bots, including a reported 16.7% spoof rate for requests claiming `ChatGPT-User` in its dataset. That is not a correction factor for your logs. Verify claimed requests against the providers’ published IP ranges, including OpenAI’s `chatgpt-user.json` and Perplexity’s `perplexity-user.json`, before counting them. Put unverified claims in a separate bucket rather than calling them assistant visits.

### Choose four URL cohorts

Pick **five to ten URLs per cohort** before the clock starts. This prevents a busy documentation page from hiding a silent commercial section:

- **Transactional money pages:** Pricing, enterprise demo, and high-value service pages.
- **Category and solution hubs:** Broad pages such as `/solutions/b2b-ecommerce` or `/features/lead-routing`.
- **Editorial explainers:** Narrow articles on mechanics, workflows, or technical comparisons.
- **Technical references and changelogs:** API pages, setup instructions, and dated updates.

Record each URL as it exists on Day 0. If documentation gets requests and pricing does not, you have a cohort-level question worth investigating. You do not yet have proof that the pricing copy is bad.

### Write 15 buyer prompts

Choose 15 questions a qualified prospect might ask while considering a purchase. Skip branded vanity prompts and vague queries such as “Best PPC tools.” Use functional constraints instead: _“Which PPC management platforms automate bid adjustments continuously without charging a percentage of ad spend?”_ or _“How do I configure server-side offline conversion tracking for Shopify on Google Ads without third-party cookie loss?”_

If you run paid search, take your top 20 converting exact-match search terms from the Google Ads search terms report and use them to help draft the 15 questions. **Freeze the prompt list for the baseline.** Changing questions halfway through changes the test.

![Diagram showing a web firewall blocking a headless AI bot while allowing a browser request](https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/09c166a5-6346-42ec-805c-2f094cc89db7.png)

### Check the edge before starting the clock

Inspect WAF events, bot rules, and your served `robots.txt`. Check whether managed AI-bot settings alter access or issue challenges; [practitioners auditing edge bot blocking](https://www.reddit.com/r/SEO/comments/1srlq0o/cloudflare_has_been_quietly_blocking_gptbot_and/) have raised this as a source of confusing reports. Do not assume a page loading in your browser proves a headless fetcher can load it too.

Record blocks and challenges on Day 0. If a verified request gets a 403 at the edge, your origin will not log a successful visit. That is an access problem to investigate before any copy test. Do not broadly disable your WAF to make a chart look better.

## Days 1–30: Collect three streams without changing the site

Run the baseline for **30 consecutive calendar days**. Avoid redesigns, URL migrations, and major copy changes during that window. Keep a dated note of unavoidable changes so you do not later credit a headline for a firewall adjustment.

1. **Verified fetches and responses.** For each URL, count verified GET requests from `ChatGPT-User` and `Perplexity-User`. Record status codes, edge decisions, and timing where your logs expose it. Look for errors, challenges, timeouts, and slow responses rather than treating every `200 OK` as a successful retrieval. [This discussion of Perplexity retrieval and latency](https://www.aigcmkt.com/en/perplexitybot-configuration.html) describes tight time budgets, but your logs cannot tell you the precise point at which an assistant abandoned a response.
2. **What the initial response contains.** Fetch the URL as a bot would receive it and inspect the returned HTML, not just the page after your browser runs JavaScript. [WISLR’s analysis of 12,099 AI bot requests](https://www.wislr.com/articles/ai-bot-behavior-log-analysis/) describes `ChatGPT-User` requesting raw HTML without accompanying CSS, JavaScript bundles, or images. If product specifications or pricing appear only after client-side hydration, a successful status code may still deliver too little text to use.
3. **Prompt answers and citations.** Run the scheduled prompts, record what appears, and compare submission times with your request logs. This stream tells you what your own checks produced. It does not turn every coincident bot visit into a proven response to your prompt.

![Server access log entries beside a browser displaying an AI assistant’s cited answer](https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/3cffed24-bf21-4547-acda-6fc832ebab2d.png)

### Run the same prompts every seven days

At seven-day intervals, submit the 15 prompts to ChatGPT Search and Perplexity Pro in clean browser sessions with cleared cookies and history. For each answer, record:

- **Your presence:** A clickable citation, an unlinked mention, or no appearance.
- **Other sources:** A direct competitor, directory, or independent review site cited instead.
- **Nearby requests:** Whether either verified user-action bot requested a cohort URL within three minutes of submission, and what the edge and origin returned.

Four rounds of 15 prompts across two engines give you **120 prompt runs**. Keep the exact prompts and procedure stable. A request close to submission time is useful evidence, but not a guarantee that the assistant used your page. A citation without a matching live request is possible too; the answer may draw on cached index material.

## Day 30: Sort the results by failure mode

Compare verified requests, response quality, delivered HTML, and prompt answers. Use these as **diagnostic buckets, not automatic verdicts**. Some URLs will not fit neatly into one.

### 1. No verified user-action fetches

Stop rewriting prose long enough to check the route to the page. Did verified requests hit a WAF rule? Did the origin receive search-indexer traffic? Are canonical tags and indexable URLs consistent? Do the selected prompts give this page a plausible reason to appear?

Zero live requests during your test means you did not observe a user-action fetch for that URL. It does **not** prove the page is absent from an index: the assistant may not have browsed, may have used cached information, or may have chosen other candidates. If you find edge blocks, fix access first. If you find no blocks, investigate indexability and the page’s fit for the prompts before commissioning a rewrite.

### 2. Verified fetches, no citations

Now inspect what the bot received. A fetch followed by no citation makes passage extraction worth testing, particularly when the initial HTML omits the answer or splits its constraints across an accordion, several sections, and a client-rendered table. A concise, self-contained answer in plain HTML gives the assistant something it can use without assembling your pitch from spare parts.

Do not call extraction the proven cause solely because a fetch occurred. The assistant can fetch a readable page and still choose a different source. Compare the delivered passage with the exact prompt and the cited alternatives. **Fix missing or fragmented answers before polishing adjectives.**

### 3. Citations go to the wrong page

Suppose a technical note gets the fetches while your commercial hub sits untouched. That pattern is worth keeping, not deleting. In [our look at a changelog attracting live visits while category pages did not](https://groas.com/post/178-chatgpt-visits-to-one-changelog-zero), the useful technical page provided a route into a commercial section that was otherwise missing the action.

Preserve the page assistants find. Add clear internal links to the relevant money page and make the commercial destination explain the same subject concretely. The practical question is not “How do I force a pricing-page citation?” It is “Where should a qualified reader go after the page that earned the visit?”

## Follow-up: Change one lever per cohort

Use the first 30 days as your baseline, then give each affected cohort a defined change for the next test window. Do not overhaul the whole site and call the resulting movement a copywriting win.

- **No-fetch cohort:** Check indexing, canonicals, and edge delivery. Where verified bot requests are incorrectly challenged, adjust the relevant rule for the verified ranges. Work on response time if logs show slow delivery. Leave body copy alone while you test access.
- **Fetched, uncited cohort:** Keep the URL and metadata stable. Put the answer and its important qualifications together in a standalone block of roughly 150–250 words, and make sure essential facts appear in the initial HTML. Then see whether the fetch-and-citation pattern changes.
- **Wrong-page cohort:** Keep the technical page live. Add contextual links and explicit commercial anchor text toward the appropriate money page; improve the hub’s factual structure without erasing what made the technical page useful.

**Change one class of problem at a time.** Otherwise, even a better result leaves you guessing which edit mattered.

## Limits: A log is evidence, not an assistant transcript

Three blind spots belong beside every result:

- **Model memory:** An assistant can answer without a live web lookup. Your logs measure requests to your site, not what the model remembers.
- **Cached material:** A citation may appear without a new GET to your server if the assistant uses indexed or cached content.
- **Small samples:** A low-traffic domain may see only a handful of verified user-action fetches in 30 days. Extend collection to 60 or 90 days if the baseline is too thin to compare, and keep manual checks focused on your commercial questions.

Do not mistake a zero in a narrow sample for a permanent verdict. Do not mistake a `200 OK` for a readable page. Both errors send teams back to rewriting copy because it is the lever they can see.

## One-page diagnostic scorecard

| Observation                          | Investigate first                                                            | Next controlled change                                        |
| :----------------------------------- | :--------------------------------------------------------------------------- | :------------------------------------------------------------ |
| **No verified user-action fetches**  | Edge blocks, indexability, and whether the prompts surface the URL           | Fix confirmed access or indexing problems before editing copy |
| **Verified fetches, no citations**   | Response quality, initial HTML, and whether a usable answer appears together | Test a self-contained answer in server-delivered HTML         |
| **Citations favor a technical page** | What that page answers and where it sends the reader                         | Keep it live; link it to a stronger commercial destination    |

I have spent enough time watching teams debate wording while a more basic part of the system was broken. Search has not cured that habit. After 30 days, the useful result is not a prettier citation dashboard: it is knowing whether to fix access, fix the delivered answer, or keep the page that wins and improve where it leads. If the logs show no live request, start upstream. If the bot gets the page but your answer is unusable, edit the page. Let your competitors begin with the adjectives.

## Frequently Asked Questions

### Should I rewrite my pages for AI citations before checking server logs?

No. Check whether ChatGPT or Perplexity actually fetch your pages first, using server and edge logs. If a verified request never reaches a page, sharper copy will not help because the failure is upstream of the content.

### What tools should I use to see whether AI assistants fetch my pages?

Use raw NGINX or Apache access logs, or request logs from an edge provider such as Cloudflare, Fastly, or AWS CloudFront. Client-side analytics like Google Analytics or Plausible depend on scripts a retrieval agent may never run.

### Why should I check edge logs separately from origin logs?

A request blocked at the edge never appears in an origin log. If you inspect only NGINX while a WAF rejects the bot upstream, you may wrongly conclude a page got zero fetches when the bot did try to visit.

### Which user agents matter when testing live AI citations?

Inspect verified GET requests from ChatGPT-User and Perplexity-User, which act on live user requests. Training crawlers like GPTBot and ClaudeBot crawl asynchronously and do not indicate someone asked about your product, while OAI-SearchBot and PerplexityBot are search indexers to track separately.

### Can I trust the User-Agent string when counting AI bot requests?

No, a User-Agent is a label anyone can print. Verify claimed requests against the providers' published IP ranges, such as OpenAI's chatgpt-user.json and Perplexity's perplexity-user.json, and put unverified claims in a separate bucket rather than counting them as assistant visits.

### How do I set up the 30-day log test for AI citations?

Pick five to ten URLs in four cohorts (transactional money pages, category hubs, editorial explainers, and technical references), write 15 buyer prompts using functional constraints, and freeze the list for 30 consecutive days. Collect verified fetches, delivered HTML, and prompt answers every seven days across ChatGPT Search and Perplexity Pro.

### What should I do if a page gets AI bot fetches but no citations?

Inspect what the bot received. If the initial HTML omits the answer or splits constraints across accordions and client-rendered tables, put a concise, self-contained answer of roughly 150-250 words in a standalone block that appears in the initial HTML, then see whether the fetch-and-citation pattern changes.

### Does a citation mean the AI assistant fetched my page live?

Not necessarily. An assistant can answer from model memory or cached index material without a new GET to your server, so a citation without a matching live request is possible. Logs measure requests to your site, not what the model remembers.

## Pay For Results, Not For Hours

Businesses buy the outcome, agencies resell it, and groas answers for it either way.

[See If You Qualify](https://groas.typeform.com/to/xC1bQNUT)

## Structured data

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://graph.groas.com/entity/groas","name":"groas","url":"https://groas.com/","logo":{"@type":"ImageObject","url":"https://groas.com/icon-512.png","width":512,"height":512},"sameAs":["https://www.linkedin.com/company/groas/"]},{"@type":"WebSite","@id":"https://groas.com/#website","name":"groas","url":"https://groas.com/","inLanguage":"en-US","publisher":{"@id":"https://graph.groas.com/entity/groas"}},{"@type":"WebPage","@id":"https://groas.com/post/are-ai-assistants-even-reading-your-page#webpage","url":"https://groas.com/post/are-ai-assistants-even-reading-your-page","name":"Before You Rewrite for AI Citations, Run This 30-Day Server Log Test","description":"Use a 30-day server log protocol to see whether ChatGPT and Perplexity fetch your pages, what they receive, and which fix to test before rewriting copy.","inLanguage":"en-US","isPartOf":{"@id":"https://groas.com/#website"},"about":{"@id":"https://graph.groas.com/entity/groas"},"primaryImageOfPage":{"@type":"ImageObject","url":"https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/5dd082f1-dd0f-4f61-9b46-8931ae33f61e.png"},"breadcrumb":{"@id":"https://groas.com/post/are-ai-assistants-even-reading-your-page#breadcrumb"},"dateModified":"2026-10-11T05:38:28.597Z","mainEntity":{"@id":"https://groas.com/post/are-ai-assistants-even-reading-your-page#article"}},{"@type":"BreadcrumbList","@id":"https://groas.com/post/are-ai-assistants-even-reading-your-page#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://groas.com/"},{"@type":"ListItem","position":2,"name":"Blog","item":"https://groas.com/blog"},{"@type":"ListItem","position":3,"name":"Before You Rewrite for AI Citations, Run This 30-Day Server Log Test","item":"https://groas.com/post/are-ai-assistants-even-reading-your-page"}]},{"@type":"BlogPosting","@id":"https://groas.com/post/are-ai-assistants-even-reading-your-page#article","url":"https://groas.com/post/are-ai-assistants-even-reading-your-page","mainEntityOfPage":{"@id":"https://groas.com/post/are-ai-assistants-even-reading-your-page#webpage"},"isPartOf":{"@id":"https://groas.com/blog#blog"},"headline":"Before You Rewrite for AI Citations, Run This 30-Day Server Log Test","description":"Use a 30-day server log protocol to see whether ChatGPT and Perplexity fetch your pages, what they receive, and which fix to test before rewriting copy.","datePublished":"2026-10-11T05:38:28.522Z","dateModified":"2026-10-11T05:38:28.597Z","image":{"@type":"ImageObject","url":"https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/5dd082f1-dd0f-4f61-9b46-8931ae33f61e.png"},"author":{"@id":"https://groas.com/author/alexander-perelman#person"},"publisher":{"@id":"https://graph.groas.com/entity/groas"},"keywords":"test, before, page, run, log, 30-day server, server log, server","wordCount":2115,"inLanguage":"en-US"},{"@type":"Person","@id":"https://groas.com/author/alexander-perelman#person","name":"Alexander Perelman","jobTitle":"Head Of Product @ groas","description":"Ex Goldman Sachs and Ex Stanford Computer Science","image":"https://groas.com/media/blog/4f17cac3a81acc224cc1b2cfab8f94bbc295a3089d13e8421125771ef8d4064f.jpg","worksFor":{"@id":"https://graph.groas.com/entity/groas"},"sameAs":["https://www.linkedin.com/in/alexander-433793253/"]},{"@type":"FAQPage","@id":"https://groas.com/post/are-ai-assistants-even-reading-your-page#faq","mainEntity":[{"@type":"Question","name":"Should I rewrite my pages for AI citations before checking server logs?","acceptedAnswer":{"@type":"Answer","text":"No. Check whether ChatGPT or Perplexity actually fetch your pages first, using server and edge logs. If a verified request never reaches a page, sharper copy will not help because the failure is upstream of the content."}},{"@type":"Question","name":"What tools should I use to see whether AI assistants fetch my pages?","acceptedAnswer":{"@type":"Answer","text":"Use raw NGINX or Apache access logs, or request logs from an edge provider such as Cloudflare, Fastly, or AWS CloudFront. Client-side analytics like Google Analytics or Plausible depend on scripts a retrieval agent may never run."}},{"@type":"Question","name":"Why should I check edge logs separately from origin logs?","acceptedAnswer":{"@type":"Answer","text":"A request blocked at the edge never appears in an origin log. If you inspect only NGINX while a WAF rejects the bot upstream, you may wrongly conclude a page got zero fetches when the bot did try to visit."}},{"@type":"Question","name":"Which user agents matter when testing live AI citations?","acceptedAnswer":{"@type":"Answer","text":"Inspect verified GET requests from ChatGPT-User and Perplexity-User, which act on live user requests. Training crawlers like GPTBot and ClaudeBot crawl asynchronously and do not indicate someone asked about your product, while OAI-SearchBot and PerplexityBot are search indexers to track separately."}},{"@type":"Question","name":"Can I trust the User-Agent string when counting AI bot requests?","acceptedAnswer":{"@type":"Answer","text":"No, a User-Agent is a label anyone can print. Verify claimed requests against the providers' published IP ranges, such as OpenAI's chatgpt-user.json and Perplexity's perplexity-user.json, and put unverified claims in a separate bucket rather than counting them as assistant visits."}},{"@type":"Question","name":"How do I set up the 30-day log test for AI citations?","acceptedAnswer":{"@type":"Answer","text":"Pick five to ten URLs in four cohorts (transactional money pages, category hubs, editorial explainers, and technical references), write 15 buyer prompts using functional constraints, and freeze the list for 30 consecutive days. Collect verified fetches, delivered HTML, and prompt answers every seven days across ChatGPT Search and Perplexity Pro."}},{"@type":"Question","name":"What should I do if a page gets AI bot fetches but no citations?","acceptedAnswer":{"@type":"Answer","text":"Inspect what the bot received. If the initial HTML omits the answer or splits constraints across accordions and client-rendered tables, put a concise, self-contained answer of roughly 150-250 words in a standalone block that appears in the initial HTML, then see whether the fetch-and-citation pattern changes."}},{"@type":"Question","name":"Does a citation mean the AI assistant fetched my page live?","acceptedAnswer":{"@type":"Answer","text":"Not necessarily. An assistant can answer from model memory or cached index material without a new GET to your server, so a citation without a matching live request is possible. Logs measure requests to your site, not what the model remembers."}}]}]}
```
