On our two busiest paths, only 19–24% of requests claiming to be AI bots passed verification. Three category pages drew thousands of claimed hits each but no verified answer-time fetches. Two how-to posts drew 135 and 144. If your AI visibility dashboard leads with crawler visits, it leads with the least useful number in our logs.
I used to tell clients to watch crawl volume as an early signal for AI search. I was wrong to give it that much weight. For 30 days, we measured requests to our own site at the edge and separated bot claims from verified bots, then separated general crawling from user-triggered fetches. The columns told different stories.
A failed verification check does not prove who sent a request. It does mean we cannot responsibly count that request as a visit from the provider it named. And a live fetch does not prove a page was cited in an answer. Those limits matter. Even with them, the decision is clear: measure verified answer-time fetches by page and assistant, then use prompt checks to investigate citations. Do not treat a pile of bot-labelled hits as visibility.
The crawler count starts with a self-declared name
A request can call itself GPTBot by putting User-Agent: GPTBot/1.4 in a header. That string is trivial to fake. Before you count the request as a provider visit, you need to check the sending IP against that provider’s published ranges or use a forward-confirmed reverse-DNS check where appropriate. A user-agent match alone verifies nothing.
The distinction is not academic. In a two-week analysis, HUMAN Security found that about one in 17 to 18 requests claiming an AI crawler identity was spoofed. That is a finding from its analysis, not an estimate of the spoofed share on our site. Our logs establish a narrower result: most bot-labelled requests on our busiest paths did not pass our provider checks. We cannot assign every failed request a motive or an operator.
The label also hides different kinds of legitimate activity. OpenAI publishes separate identities and IP files for its crawlers. GPTBot is associated with training; ChatGPT-User makes user-triggered fetches. Their roles and robots.txt directives differ. Adding their hits together gives you a total, but not an answer to a useful question.
Practical takeaway: verify the visitor first. Then ask what job that visitor was doing.
What our 30-day edge log can actually show
We logged requests at the CDN edge for 30 days on our own domain. For this analysis, we grouped requests by exact path and used three measures:
| Measure | What we counted | What it does not establish |
|---|---|---|
| Bot-claimed hits | Requests with an AI bot name in the user-agent | That the named provider sent them |
| Verified hits | Requests that passed our provider IP-range or forward-confirmed reverse-DNS check | That the provider used the page in an answer |
| Verified answer-time fetches | Verified ChatGPT-User or Claude-User requests for a page | That the page appeared as a citation |
For OpenAI, we checked sending IPs against its published CIDRs; we used the available provider file for Anthropic and forward-confirmed reverse DNS where ranges were not published. These are our edge-log counts, not a cross-site benchmark. The logs tell us which paths were requested and by which verified agent. They do not show the answer a user saw or why an assistant chose a URL.
That last boundary is easy to miss because prompt tracking and crawler-log tracking measure different things. A bot can fetch a page without citing it. An answer can cite a page without fetching it during the window you are watching. Whenever I say a page had an answer-time fetch below, I mean exactly that, not a confirmed citation.
There is a cost reason to keep an eye on crawling too. Read the Docs reported that blocking AI crawlers reduced its daily traffic from 800GB to 200GB. That example makes resource use worth measuring; it does not make crawl volume a visibility metric. A separate proxy study across 69 sites checked OpenAI requests against published CIDRs rather than trusting user-agent strings. We applied that same basic discipline to a smaller scope: one domain, one month.
Finding 1: The busiest paths failed the identity check most often
/robots.txt drew 13,583 bot-claimed hits, of which 3,197 passed verification. The homepage drew 11,481 claimed hits, of which 2,173 passed. Both sets of figures come from our 30-day edge log.
| Path | Bot-claimed hits | Verified hits | Verified share |
|---|---|---|---|
/robots.txt | 13,583 | 3,197 | 23.5% |
/ (homepage) | 11,481 | 2,173 | 18.9% |
On these two paths, the claimed count was roughly four to five times the verified count. That does not mean we proved every other request was fake. Some may have failed a check for reasons these figures cannot resolve. It does mean a report that calls all 25,064 claimed hits “AI crawler visits” overstates what we can attribute to verified providers.
The paths themselves help explain why the discrepancy matters. robots.txt and the homepage attract discovery requests as well as whatever else arrives wearing a bot label. They are sensible places for a crawler to start, but they are poor proxies for whether a particular page helped answer a buyer’s question. If the first chart in your report ranks those paths as AI-search winners, the chart needs a different job description.
Practical takeaway: keep claimed hits for debugging. Use verified hits when you report provider activity, and do not call either number citations.
Finding 2: Thousands of category-page hits led to no recorded answer-time fetches
Three category pages each drew roughly 3,300–3,750 bot-claimed hits during the same 30 days. Their verified shares were higher than the homepage’s, but we recorded zero verified ChatGPT-User or Claude-User fetches for those pages. Two how-to posts, neither in the top ten by crawl volume, recorded 135 and 144 answer-time fetches respectively. Again, these are our edge-log results, not a claim about what appeared in every live answer.
| Page group | Bot-claimed hits in our log | Verified answer-time fetches |
|---|---|---|
| Three category pages | About 3,300–3,750 each | 0 each |
| How-to post 1 | Outside the top ten by crawl volume | 135 |
| How-to post 2 | Outside the top ten by crawl volume | 144 |
The high-volume pages were useful to crawlers as entry points or taxonomy. The two guides offered more specific material for a user-triggered fetch. That is an interpretation of the pattern, not proof of an assistant’s reasoning. The observation we can defend is simpler: ranking pages by claimed crawl volume would have hidden both posts with answer-time activity.
Broader figures show why a large crawl count deserves skepticism, though they cannot explain our individual pages. A Cloudflare-based account reports about 1,255 GPTBot fetches per referral and about 20,000 ClaudeBot fetches per referral. Those are aggregate ratios, not conversion rates you can apply to a guide on our site. The useful point is the gap between being fetched and sending a referral. Even the 10,000-crawls-to-100-referrals worked example illustrates a different measurement from citations or answer-time fetches.
A category page with zero recorded live fetches might still appear in an answer through a cached source. Our log cannot rule that out. It can tell us where a fresh, verified user-triggered request did—and did not—land during this window. That is enough to change which pages I would inspect first.

Finding 3: ChatGPT and Claude did not pull the same pages
The assistant split matters as much as the page split. In our log, ChatGPT-User fetched the homepage at answer time and repeatedly pulled one how-to guide. Claude-User largely passed over the homepage and pulled a different guide, the AI Max walkthrough. Together, those two posts account for the 135 and 144 answer-time fetches above. The category pages recorded none.
I would not infer a permanent preference from one month on one site. Nor can the logs tell us whether indexing, freshness, formatting, or the questions being asked caused the difference. They show that a single combined “AI bot” total would conceal it. Separate resampling work found limited overlap in domains cited by ChatGPT and Perplexity; that is citation evidence from a different comparison, not validation of a cause in our logs.
Practical takeaway: break out verified answer-time fetches by page and assistant. If one guide draws ChatGPT-User requests and another draws Claude-User requests, a blended total is not a plan for either page.
Run the four-column test before rewriting anything
Here is the question I would put at the top of the lab notebook: which paths attract bot claims, which requests pass verification, and which pages get fetched at answer time? Run the check before changing the pages so you have a baseline to compare with the next window. This is a prospective test for your domain, not a claim that your distribution will match ours.

- Set the window and preserve the raw fields. Use a 30-day log window, with a one-week baseline you can inspect before making changes. Keep timestamp, exact path, user-agent and source IP from CDN or origin logs. Do not roll every URL into a domain-wide total.
- Verify before classifying. Match claimed provider requests against the relevant published IP lists, and use a forward-confirmed reverse-DNS check where appropriate. Keep failed or unresolved claims separate rather than silently counting them as provider traffic. The curl and CDN audit is a companion walkthrough for the checks.
- Build four columns for each path. Record claimed hits, verified hits, verified ChatGPT-User and Claude-User fetches, and citations observed in a weekly prompt check. Keep citations in their own column: the logs alone cannot produce that count.
- Hold the comparison steady. Use the same prompt set and engines for each weekly check. Compare like periods after a page change; do not credit a rewrite because one screenshot moved.
My expectation, based on our site, is that claimed hits will substantially exceed verified hits on /robots.txt and /, while verified answer-time requests will concentrate on a few specific guides. The mechanism is the separation between broad crawling, unverified bot claims and requests triggered while an assistant is serving a user. But your site may reverse the page pattern. If category pages get the live fetches, work from that list rather than ours. If claimed and verified counts are close, you have learned that the first filter matters less for your traffic than it did for ours.
Either result is useful. Let your verified per-page requests determine where you investigate next, and let the independent prompt checks test whether that attention corresponds to visible answers.
Prompt trackers belong beside the logs, not in place of them
A prompt tracker asks whether your brand or page appeared for a set of questions. An edge log asks whether a verified agent requested a URL. Neither can substitute for the other. A citation may come from an earlier fetch; a fresh fetch may never become a citation. That is why I would not label the 135 and 144 requests in our log “135 and 144 citations.”
The tracker has its own noise. In the identical-prompt resampling analysis, resampling accounted for 34.8% of measured movement, while reported weekly citation churn reached 56% in Google AI Mode and 74% in ChatGPT. Those figures describe that analysis, not the expected wobble of every prompt set. They are enough to make a Tuesday screenshot a poor daily scorecard.
Use a fixed set of buyer questions across multiple engines, check it weekly, and read the multiweek pattern rather than the single result. The 50-plus-prompts, three-plus-engines approach gives you a way to watch where mentions or citations appear. Then return to the logs to see whether verified answer-time fetches concentrate on particular URLs. If the tracker and logs disagree, investigate the disagreement; do not force them into one invented visibility score.
Practical takeaway: prompt checks show observed answer appearances. Verified logs show fresh requests. Keep both labels intact.
The decision these numbers support
For the next reporting cycle, I would change three things:
- Put verified answer-time fetches by page and assistant on the main report. Leave claimed hits in the raw log. They help diagnose traffic, but they should not sit on a slide labelled AI visibility.
- Keep training and user-triggered activity separate. GPTBot and ChatGPT-User do different jobs, and their robots.txt controls differ. Decide which activity you intend to allow, then measure it under its own name.
- Inspect the pages with recorded live-fetch activity before chasing crawl leaders. On our site, that means the two how-to posts. I would check whether each gives a direct answer near the top, presents its facts clearly in clean HTML, and has useful links from the heavily crawled category pages. Then I would watch the same per-URL measures over the next four weeks. A rewrite is a test, not a guaranteed way to earn a citation.
Stop asking how many times something calling itself an AI bot visited. On our site, the busiest paths passed verification only 19–24% of the time, and three heavily crawled category pages recorded no verified answer-time fetches. Two guides recorded 135 and 144. The decision is to measure that split on your own domain before choosing a page to rewrite. If your numbers point somewhere else, follow them.
Frequently asked questions
What percentage of AI bot requests actually passed verification on the busiest site paths?
On the two busiest paths, only 19 to 24 percent of requests claiming to be AI bots passed verification over 30 days of edge logs. The homepage drew 11,481 claimed hits with 2,173 verified, and robots.txt drew 13,583 claimed hits with 3,197 verified.
Can I trust a request that identifies itself as GPTBot in its user-agent?
No. A request can put GPTBot/1.4 in the User-Agent header, and that string is trivial to fake. Before counting a request as a provider visit, check the sending IP against the provider's published ranges or use a forward-confirmed reverse-DNS check where appropriate.
What is the difference between bot-claimed hits, verified hits, and verified answer-time fetches?
Bot-claimed hits are requests with an AI bot name in the user-agent, which does not prove the named provider sent them. Verified hits passed a provider IP-range or forward-confirmed reverse-DNS check but do not prove the page was used in an answer. Verified answer-time fetches are verified ChatGPT-User or Claude-User requests for a page, which still do not prove the page appeared as a citation.
Did the most crawled pages get the most answer-time fetches?
No. Three category pages each drew roughly 3,300 to 3,750 bot-claimed hits but recorded zero verified ChatGPT-User or Claude-User fetches. Two how-to posts, neither in the top ten by crawl volume, recorded 135 and 144 answer-time fetches respectively.
Do ChatGPT and Claude fetch the same pages on a site?
In the 30-day log described, they did not. ChatGPT-User fetched the homepage at answer time and repeatedly pulled one how-to guide, while Claude-User largely passed over the homepage and pulled a different guide, the AI Max walkthrough. A single combined AI bot total would conceal that difference.
Are GPTBot and ChatGPT-User the same crawler?
No. GPTBot is associated with training, while ChatGPT-User makes user-triggered fetches, and their roles and robots.txt directives differ. OpenAI publishes separate identities and IP files for its crawlers, so adding their hits together gives a total but not an answer to a useful question.
How should I measure AI bot activity on my own site before rewriting pages?
Use a 30-day log window with a one-week baseline, keeping timestamp, exact path, user-agent and source IP. Verify claimed provider requests against published IP lists, then record four columns per path: claimed hits, verified hits, verified ChatGPT-User and Claude-User fetches, and citations from a weekly prompt check with a fixed prompt set. Compare like periods before crediting any page change.




