Of the 13,716 requests claiming to be AI crawlers that hit /robots.txt on groas.com in 30 days, only 3,224 were verified. The other 10,492 were noise as far as this measurement is concerned: requests that presented a crawler identity we could not validate.
That gap is why I do not start an AI visibility report with bot traffic. I start with verified live-retrieval fetches per page. In the same log window, our category archives drew thousands of verified index crawls and no verified live-retrieval fetches. Two dated posts drew 157 and 149 live-retrieval fetches each. Raw activity pointed in one direction; the requests closest to a live answer pointed in another.
First, remove the bot hits that cannot prove who sent them
A User-Agent is self-reported text. A script can send User-Agent: ChatGPT-User or User-Agent: ClaudeBot, and a dashboard that groups requests by that string will count an AI visit. Research on crawler identity has found that up to 98% of requests claiming certain AI crawler identities in raw logs can be fake. Scrapers also disguise themselves as AI assistants to get past rules that treat familiar crawler names more generously.
Here is what our CDN edge logs showed for one path:
/robots.txt, 30-day window | Requests | Share of claimed hits |
|---|---|---|
| Claimed AI crawler requests | 13,716 | 100% |
| Verified requests | 3,224 | 23.5% |
| Unverified requests | 10,492 | 76.5% |
Source: groas.com CDN edge logs. These figures describe requests to /robots.txt, not all paths on the site.
Unverified does not tell me exactly who sent a request. It tells me not to credit that request to the AI operator named in its User-Agent. The distinction matters: assigning every suspicious hit to a scraper, a probe, or a data harvester would claim more than an IP check can establish.
Basic User-Agent matching skips that check. Reverse DNS is not a dependable shortcut either; bot authentication can be difficult when requests use changing infrastructure. For this analysis, a claimed identity needed to match the relevant vendor-published IP ranges before it counted as verified. Otherwise, a chart of “AI crawler growth” could rise while genuine crawler activity stayed flat.
Three kinds of requests, three different questions
Even after verification, I would not put every bot request in one bucket. OpenAI distinguishes GPTBot, OAI-SearchBot, and ChatGPT-User; Anthropic distinguishes ClaudeBot, Claude-SearchBot, and Claude-User. Those identities point to different jobs.
| Tier | What the log shows | What it can tell you |
|---|---|---|
| Claimed hits | A request carries an AI crawler User-Agent. | Someone or something claimed that identity. |
| Verified background crawls | A validated training or search-index crawler requests a URL. | The operator fetched the page for a background task. |
| Verified live-retrieval fetches | A validated user-retrieval agent requests a URL. | The operator fetched the page in its live-retrieval workflow. |
The third tier is the useful one for this article, but it needs a careful name. A verified ChatGPT-User or Claude-User request is not proof that the page was quoted, cited, or shown to a buyer. The edge log records the fetch, not the finished answer. It is still much closer to answer-time activity than an unverified hit or a background index crawl. Our step-by-step CDN audit walks through the same separation.

Across all 114 published URLs in our 30-day rollup, the logs contained 68,410 requests claiming an AI origin. We classified 14,288 as verified indexing or training requests, or 20.9% of that raw volume. Those are different counts for different jobs, not a conversion funnel from crawl to answer. In particular, the indexing figure is not a count of every verified request on the domain.
The page table changes the story
The homepage drew substantial live-retrieval activity: 979 verified fetches, including 938 from ChatGPT-User. That is a reasonable place for an agent to look for the brand’s basic description. The more interesting result appears below it. The blog index and category archives attracted heavy background crawling, while two specific posts attracted repeated live-retrieval fetches.
| Page path or template | Verified index crawls | Verified live-retrieval fetches |
|---|---|---|
Homepage (/) | 3,112 | 979 |
Setup and data guide (/post/...) | 482 | 157 |
Changelog and updates (/post/...) | 319 | 149 |
Blog index (/blog) | 4,220 | 2 |
| Category archives (six taxonomy hubs) | 3,890 | 0 |
Source: groas.com CDN edge logs, grouped by URL path over the same 30-day window. The table shows selected paths and templates, not an exhaustive domain total. A live-retrieval fetch is a request by a verified retrieval agent; it does not establish that the URL appeared in the final answer.
The category archives received 3,890 verified index crawls and zero verified live-retrieval fetches. The blog index received 4,220 index crawls and just two live-retrieval fetches. The setup guide and changelog together received fewer index crawls than either archive group, yet each drew more than a hundred live-retrieval fetches.
I would not call that proof that category pages have no value. They can still help people navigate the site, and the logs show crawlers did request them. But if the job is to find pages a live-retrieval agent fetches, index crawl volume is a poor substitute for the per-page count. That is the measurement error I would fix before rewriting a single headline.

The robots.txt wrinkle: retrieval agents also fetch rules
There is another reason to group by path. Our logs recorded 1,194 verified live-retrieval-agent requests to /robots.txt, almost all carrying the Claude-User identity. A rules-file fetch is not an article fetch. Counting both under a single “live AI visits” total would make content look busier than it was.
The distinction also matters when setting crawler permissions. OpenAI describes how ChatGPT-User handles user-initiated requests, while operators’ bot identities and rules differ. A blanket AI-bot block can affect more than background crawling. If you intend to block training while allowing live retrieval, name the agents you mean rather than treating every AI User-Agent as the same visitor. Our AI crawler swipe file lays out configurations for that separation.
The practical logging rule is simpler: keep /robots.txt in its own row. Its requests tell you something about access checks. They do not tell you which content page fed an answer.
What the two posts had that the archives did not
The setup guide logged 157 verified live-retrieval fetches; the changelog logged 149. Both were dated and specific. They carried publication or modification timestamps, version details, and concrete configuration information. An archive offers links and short descriptions instead. If a retrieval system needs a particular operational detail, a page containing that detail is a more direct place to fetch it.
That is a mechanism, not a controlled experiment. These logs cannot isolate whether timestamps, page structure, the subject of the query, or some other factor caused the difference. The result is narrower and more useful: on this site, in this window, specific posts drew live fetches and category templates did not.
That observation changes what I would maintain first. I used to treat elaborate category structures as a default content investment. For conversational retrieval, I would put the first editing hour into the pages already being fetched: make sure the setup steps still describe the current process, the changelog is current, and the operational claims say what they mean. Do not let a vague marketing rewrite replace a precise answer just because it sounds smoother in a meeting.
I would not delete useful navigation to improve an AI metric. I would stop using navigation-page crawl counts to justify spending more time on those pages for AI visibility. Those are different decisions, and the table supports only the second one.
Run the same check on your site
Here is the test I would run before trusting an AI visibility dashboard. Question: Which content URLs receive verified live-retrieval fetches, and how much claimed bot activity fails identity checks?
Setup and control. Use a 30-day edge or server-log window with client IP addresses, paths, timestamps, and full User-Agent strings. Keep /robots.txt separate from document paths throughout. Without that control, rules checks can swamp the page-level result. Cloudflare Enterprise Logpush, AWS CloudFront, Fastly, or origin access logs from Nginx or Caddy can provide the request data if your setup records those fields.
- Collect claimed requests. Filter GET requests for the crawler identities you intend to examine, including
ChatGPT-User,GPTBot,OAI-SearchBot,Claude-User,ClaudeBot,Claude-SearchBot, andPerplexityBot. Retain the full request record; do not reduce it to a daily total yet. - Verify each identity. Compare the client IP with the vendor-published range for the claimed agent. OpenAI publishes separate lists at
openai.com/chatgpt-user.json,openai.com/searchbot.json, andopenai.com/gptbot.json; use Anthropic’s published ranges for its agents. Do not classify a request as verified just because its User-Agent or reverse DNS looks plausible. - Separate the jobs. Group validated requests by background crawler versus live-retrieval agent, then by exact URL path. Put
/robots.txtin a separate group. Keep failed identity checks visible as unverified claimed hits rather than mixing them into either verified group. - Read the page table. Compare verified index crawls and verified live-retrieval fetches for each content path. Flag pages with repeated live fetches for an accuracy review. Investigate commercially important pages with no verified activity before assuming they need more copy.

Measure two ratios, but keep their denominators in view:
- Verified share: verified requests divided by claimed requests for the same path or cohort. On our
/robots.txtcohort, that was 3,224 out of 13,716, or 23.5%. It says how much of the claimed activity passed identity checks, not how often the page appeared in answers. - Live-retrieval share: verified live-retrieval requests divided by verified requests for the same content path. Use it alongside the fetch count. A tiny page sample can produce an impressive percentage without much activity behind it.
I would expect an archive-heavy site to show more verified background crawls than live-retrieval fetches on its category pages. The mechanism is straightforward: an archive exposes links, while a focused page holds the detail a retrieval agent may need. That is an expectation for the test, not a result I can promise for your site.
A high-index, zero-live page may be doing navigation work rather than answer work. A commercially important page with neither kind of request warrants an access check before a content rewrite. A page with recurring verified live fetches deserves a close review of its dates, instructions, and claims. In every case, the logs stop at the request: they do not reveal the prompt, the final citation, or revenue.
The decision the numbers support
The 30-day result does not make raw crawl volume useless. It gives that volume a smaller job. Claimed hits help you spot traffic worth checking; verified background crawls show that an operator fetched a page for another purpose. Neither is a stand-in for answer-time retrieval.
On groas.com, the distinction was hard to miss: 76.5% of claimed AI crawler hits to /robots.txt failed verification, category archives drew 3,890 index crawls and no live-retrieval fetches, and two specific posts drew 157 and 149 live fetches. Measure verified live-retrieval fetches by content page, with rules-file requests excluded. Then spend the next editing hour where those requests actually land.

