---
title: "We Counted AI Bot Hits for 30 Days. The Useful Number Was Fetches Per Page"
description: "Of the 13,716 requests claiming to be AI crawlers that hit /robots.txt on groas.com in 30 days, only 3,224 were verified."
url: "https://groas.com/post/we-counted-every-ai-bot-hit-on-our-site"
image: "https://groas.com/media/blog/ded929a3f8f0993b8f15bd4bb371f97a69d446561de97fc98e695256dd9048f8.png"
published: "2026-10-07T05:20:02.860Z"
modified: "2026-10-07T05:20:02.860Z"
---

[AI For Google Ads](https://groas.com/category/ai-for-google-ads) · October 7, 2026 · 9 min read

# We Counted AI Bot Hits for 30 Days. The Useful Number Was Fetches Per Page

[DavidFounder & CEO @ groas](https://groas.com/author/david)

![Tabletop model town: a huge crowd, many of them paper fakes, mills outside a dark grand hall while a few real figures queue at a small glowing workshop.](https://groas.com/media/blog/ded929a3f8f0993b8f15bd4bb371f97a69d446561de97fc98e695256dd9048f8.png)

Of the 13,716 requests claiming to be AI crawlers that hit `/robots.txt` on groas.com in 30 days, **only 3,224 were verified**. The other 10,492 were noise as far as this measurement is concerned: requests that presented a crawler identity we could not validate.

That gap is why I do not start an AI visibility report with bot traffic. I start with **verified live-retrieval fetches per page**. In the same log window, our category archives drew thousands of verified index crawls and no verified live-retrieval fetches. Two dated posts drew 157 and 149 live-retrieval fetches each. Raw activity pointed in one direction; the requests closest to a live answer pointed in another.

## First, remove the bot hits that cannot prove who sent them

A User-Agent is self-reported text. A script can send `User-Agent: ChatGPT-User` or `User-Agent: ClaudeBot`, and a dashboard that groups requests by that string will count an AI visit. Research on crawler identity has found that [up to 98% of requests claiming certain AI crawler identities in raw logs can be fake](https://hey-eye.gr/blog/verify-real-ai-crawlers-vs-spoofed-bots/). Scrapers also [disguise themselves as AI assistants](https://fingerprint.com/blog/product-update-ai-assistant-detection/) to get past rules that treat familiar crawler names more generously.

Here is what our CDN edge logs showed for one path:

| `/robots.txt`, 30-day window | Requests | Share of claimed hits |
| ---------------------------- | -------- | --------------------- |
| Claimed AI crawler requests  | 13,716   | 100%                  |
| Verified requests            | 3,224    | 23.5%                 |
| Unverified requests          | 10,492   | 76.5%                 |

_Source: groas.com CDN edge logs. These figures describe requests to `/robots.txt`, not all paths on the site._

**Unverified does not tell me exactly who sent a request.** It tells me not to credit that request to the AI operator named in its User-Agent. The distinction matters: assigning every suspicious hit to a scraper, a probe, or a data harvester would claim more than an IP check can establish.

Basic User-Agent matching skips that check. Reverse DNS is not a dependable shortcut either; [bot authentication can be difficult when requests use changing infrastructure](https://blog.cloudflare.com/web-bot-auth/). For this analysis, a claimed identity needed to match the relevant vendor-published IP ranges before it counted as verified. Otherwise, a chart of “AI crawler growth” could rise while genuine crawler activity stayed flat.

### Three kinds of requests, three different questions

Even after verification, I would not put every bot request in one bucket. [OpenAI distinguishes GPTBot, OAI-SearchBot, and ChatGPT-User](https://geoiphub.com/blog/gptbot-vs-oai-searchbot-vs-chatgpt-user/); [Anthropic distinguishes ClaudeBot, Claude-SearchBot, and Claude-User](https://www.searchenginejournal.com/anthropics-claude-bots-make-robots-txt-decisions-more-granular/568253/). Those identities point to different jobs.

| Tier                                | What the log shows                                           | What it can tell you                                          |
| ----------------------------------- | ------------------------------------------------------------ | ------------------------------------------------------------- |
| **Claimed hits**                    | A request carries an AI crawler User-Agent.                  | Someone or something claimed that identity.                   |
| **Verified background crawls**      | A validated training or search-index crawler requests a URL. | The operator fetched the page for a background task.          |
| **Verified live-retrieval fetches** | A validated user-retrieval agent requests a URL.             | The operator fetched the page in its live-retrieval workflow. |

The third tier is the useful one for this article, but it needs a careful name. A verified `ChatGPT-User` or `Claude-User` request is **not proof that the page was quoted, cited, or shown to a buyer**. The edge log records the fetch, not the finished answer. It is still much closer to answer-time activity than an unverified hit or a background index crawl. Our [step-by-step CDN audit](https://groas.com/post/can-chatgpt-actually-read-your-site-a-45) walks through the same separation.

![Three-tier sorting machine separating claimed bot hits, verified background crawls, and live-retrieval fetches](https://groas.com/media/blog/a3684b7d8398996708f60ec5e2b3fc75ac6b99a067eb289eaea236b16764ccc0.png)

Across all 114 published URLs in our 30-day rollup, the logs contained 68,410 requests claiming an AI origin. We classified 14,288 as verified indexing or training requests, or 20.9% of that raw volume. Those are **different counts for different jobs**, not a conversion funnel from crawl to answer. In particular, the indexing figure is not a count of every verified request on the domain.

#### The page table changes the story

The homepage drew substantial live-retrieval activity: 979 verified fetches, including 938 from `ChatGPT-User`. That is a reasonable place for an agent to look for the brand’s basic description. The more interesting result appears below it. The blog index and category archives attracted heavy background crawling, while two specific posts attracted repeated live-retrieval fetches.

| Page path or template                 | Verified index crawls | Verified live-retrieval fetches |
| ------------------------------------- | --------------------- | ------------------------------- |
| Homepage (`/`)                        | 3,112                 | 979                             |
| Setup and data guide (`/post/...`)    | 482                   | 157                             |
| Changelog and updates (`/post/...`)   | 319                   | 149                             |
| Blog index (`/blog`)                  | 4,220                 | 2                               |
| Category archives (six taxonomy hubs) | 3,890                 | 0                               |

_Source: groas.com CDN edge logs, grouped by URL path over the same 30-day window. The table shows selected paths and templates, not an exhaustive domain total. A live-retrieval fetch is a request by a verified retrieval agent; it does not establish that the URL appeared in the final answer._

**The category archives received 3,890 verified index crawls and zero verified live-retrieval fetches.** The blog index received 4,220 index crawls and just two live-retrieval fetches. The setup guide and changelog together received fewer index crawls than either archive group, yet each drew more than a hundred live-retrieval fetches.

I would not call that proof that category pages have no value. They can still help people navigate the site, and the logs show crawlers did request them. But if the job is to find pages a live-retrieval agent fetches, index crawl volume is a poor substitute for the per-page count. That is the measurement error I would fix before rewriting a single headline.

![Web server log pipeline separating claimed hits, background crawls, and live-retrieval requests by page](https://groas.com/media/blog/59da17d45112fa6a4ecb9e122a12c7e1c387ab37f3547cdbae9d97e2387b5661.png)

##### The robots.txt wrinkle: retrieval agents also fetch rules

There is another reason to group by path. Our logs recorded 1,194 verified live-retrieval-agent requests to `/robots.txt`, almost all carrying the `Claude-User` identity. A rules-file fetch is not an article fetch. Counting both under a single “live AI visits” total would make content look busier than it was.

The distinction also matters when setting crawler permissions. [OpenAI describes how ChatGPT-User handles user-initiated requests](https://developers.openai.com/api/docs/bots), while operators’ bot identities and rules differ. A blanket AI-bot block can affect more than background crawling. If you intend to block training while allowing live retrieval, name the agents you mean rather than treating every AI User-Agent as the same visitor. Our [AI crawler swipe file](https://groas.com/post/the-ai-crawler-swipe-file-copy-paste-rob) lays out configurations for that separation.

The practical logging rule is simpler: **keep `/robots.txt` in its own row**. Its requests tell you something about access checks. They do not tell you which content page fed an answer.

#### What the two posts had that the archives did not

The setup guide logged 157 verified live-retrieval fetches; the changelog logged 149. Both were dated and specific. They carried publication or modification timestamps, version details, and concrete configuration information. An archive offers links and short descriptions instead. If a retrieval system needs a particular operational detail, a page containing that detail is a more direct place to fetch it.

That is a mechanism, not a controlled experiment. These logs cannot isolate whether timestamps, page structure, the subject of the query, or some other factor caused the difference. The result is narrower and more useful: **on this site, in this window, specific posts drew live fetches and category templates did not**.

That observation changes what I would maintain first. I used to treat elaborate category structures as a default content investment. For conversational retrieval, I would put the first editing hour into the pages already being fetched: make sure the setup steps still describe the current process, the changelog is current, and the operational claims say what they mean. Do not let a vague marketing rewrite replace a precise answer just because it sounds smoother in a meeting.

I would not delete useful navigation to improve an AI metric. I would stop using navigation-page crawl counts to justify spending more time on those pages _for AI visibility_. Those are different decisions, and the table supports only the second one.

#### Run the same check on your site

Here is the test I would run before trusting an AI visibility dashboard. **Question:** Which content URLs receive verified live-retrieval fetches, and how much claimed bot activity fails identity checks?

**Setup and control.** Use a 30-day edge or server-log window with client IP addresses, paths, timestamps, and full User-Agent strings. Keep `/robots.txt` separate from document paths throughout. Without that control, rules checks can swamp the page-level result. Cloudflare Enterprise Logpush, AWS CloudFront, Fastly, or origin access logs from Nginx or Caddy can provide the request data if your setup records those fields.

1. **Collect claimed requests.** Filter GET requests for the crawler identities you intend to examine, including `ChatGPT-User`, `GPTBot`, `OAI-SearchBot`, `Claude-User`, `ClaudeBot`, `Claude-SearchBot`, and `PerplexityBot`. Retain the full request record; do not reduce it to a daily total yet.
2. **Verify each identity.** Compare the client IP with the vendor-published range for the claimed agent. OpenAI publishes separate lists at `openai.com/chatgpt-user.json`, `openai.com/searchbot.json`, and `openai.com/gptbot.json`; use Anthropic’s published ranges for its agents. Do not classify a request as verified just because its User-Agent or reverse DNS looks plausible.
3. **Separate the jobs.** Group validated requests by background crawler versus live-retrieval agent, then by exact URL path. Put `/robots.txt` in a separate group. Keep failed identity checks visible as unverified claimed hits rather than mixing them into either verified group.
4. **Read the page table.** Compare verified index crawls and verified live-retrieval fetches for each content path. Flag pages with repeated live fetches for an accuracy review. Investigate commercially important pages with no verified activity before assuming they need more copy.

![Three-step bot audit: collect raw logs, verify vendor IP ranges, and compare retrieval fetches by path](https://groas.com/media/blog/c6746d667ff10dafa930dd8908e26bdc540f07dd8ba80e278ee718529ca0d073.png)

**Measure two ratios, but keep their denominators in view:**

- **Verified share:** verified requests divided by claimed requests for the same path or cohort. On our `/robots.txt` cohort, that was 3,224 out of 13,716, or 23.5%. It says how much of the claimed activity passed identity checks, not how often the page appeared in answers.
- **Live-retrieval share:** verified live-retrieval requests divided by verified requests for the same content path. Use it alongside the fetch count. A tiny page sample can produce an impressive percentage without much activity behind it.

I would expect an archive-heavy site to show more verified background crawls than live-retrieval fetches on its category pages. The mechanism is straightforward: an archive exposes links, while a focused page holds the detail a retrieval agent may need. **That is an expectation for the test, not a result I can promise for your site.**

A high-index, zero-live page may be doing navigation work rather than answer work. A commercially important page with neither kind of request warrants an access check before a content rewrite. A page with recurring verified live fetches deserves a close review of its dates, instructions, and claims. In every case, the logs stop at the request: they do not reveal the prompt, the final citation, or revenue.

#### The decision the numbers support

The 30-day result does not make raw crawl volume useless. It gives that volume a smaller job. Claimed hits help you spot traffic worth checking; verified background crawls show that an operator fetched a page for another purpose. Neither is a stand-in for answer-time retrieval.

On groas.com, the distinction was hard to miss: 76.5% of claimed AI crawler hits to `/robots.txt` failed verification, category archives drew 3,890 index crawls and no live-retrieval fetches, and two specific posts drew 157 and 149 live fetches. **Measure verified live-retrieval fetches by content page, with rules-file requests excluded.** Then spend the next editing hour where those requests actually land.

## Frequently Asked Questions

### How many claimed AI crawler hits actually pass identity verification?

Of 13,716 requests claiming to be AI crawlers that hit /robots.txt on groas.com in 30 days, only 3,224 were verified, or 23.5%. The other 10,492, or 76.5%, carried a crawler identity that could not be validated and were not credited to the AI operator named in their User-Agent.

### Why can't I trust the User-Agent string in my server logs to identify AI bots?

A User-Agent is self-reported text, so any script can send User-Agent: ChatGPT-User or ClaudeBot, and a dashboard grouping by that string will count a fake visit. Research cited in the article found up to 98% of requests claiming certain AI crawler identities in raw logs can be fake. A claimed identity only counts as verified when the client IP matches the vendor-published IP ranges.

### What is the difference between a background crawl and a live-retrieval fetch?

A verified background crawl is a validated training or search-index crawler fetching a URL for a background task. A verified live-retrieval fetch is a validated user-retrieval agent, such as ChatGPT-User or Claude-User, fetching a page during its live-retrieval workflow. A live fetch still does not prove the page was quoted, cited, or shown to a buyer; the log records the fetch, not the finished answer.

### Do category pages get fetched by AI agents when they get lots of crawls?

On groas.com, category archives received 3,890 verified index crawls and zero verified live-retrieval fetches, while the blog index drew 4,220 crawls and just two live fetches. Two specific dated posts drew 157 and 149 live-retrieval fetches each. Crawl volume is therefore a poor substitute for the per-page live fetch count when judging AI visibility.

### Should robots.txt requests count as AI visits to my content?

No. The logs recorded 1,194 verified live-retrieval-agent requests to /robots.txt, and a rules-file fetch is not an article fetch. Counting both under a single live AI visits total makes content look busier than it was, so keep /robots.txt in its own row in your log analysis.

### What made the two posts attract live-retrieval fetches while the archives did not?

Both posts were dated and specific: they carried publication or modification timestamps, version details, and concrete configuration information, while an archive offers links and short descriptions. This is a mechanism, not a controlled experiment; the logs cannot isolate whether timestamps, page structure, or query subject caused the difference.

### How can I check which of my pages receive verified live-retrieval fetches?

Use a 30-day edge or server-log window with client IPs, paths, timestamps, and full User-Agent strings, and keep /robots.txt separate. Filter for claimed AI crawler identities, verify each client IP against vendor-published ranges such as openai.com/chatgpt-user.json, then compare verified index crawls and verified live-retrieval fetches by exact URL path.

## Pay For Results, Not For Hours

Businesses buy the outcome, agencies resell it, and groas answers for it either way.

[See If You Qualify](https://groas.typeform.com/to/xC1bQNUT)

## Structured data

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://graph.groas.com/entity/groas","name":"groas","url":"https://groas.com/","logo":{"@type":"ImageObject","url":"https://groas.com/icon-512.png","width":512,"height":512},"sameAs":["https://www.linkedin.com/company/groas/"]},{"@type":"WebSite","@id":"https://groas.com/#website","name":"groas","url":"https://groas.com/","inLanguage":"en-US","publisher":{"@id":"https://graph.groas.com/entity/groas"}},{"@type":"WebPage","@id":"https://groas.com/post/we-counted-every-ai-bot-hit-on-our-site#webpage","url":"https://groas.com/post/we-counted-every-ai-bot-hit-on-our-site","name":"We Counted AI Bot Hits for 30 Days. The Useful Number Was Fetches Per Page","description":"Of the 13,716 requests claiming to be AI crawlers that hit /robots.txt on groas.com in 30 days, only 3,224 were verified.","inLanguage":"en-US","isPartOf":{"@id":"https://groas.com/#website"},"about":{"@id":"https://graph.groas.com/entity/groas"},"primaryImageOfPage":{"@type":"ImageObject","url":"https://groas.com/media/blog/ded929a3f8f0993b8f15bd4bb371f97a69d446561de97fc98e695256dd9048f8.png"},"breadcrumb":{"@id":"https://groas.com/post/we-counted-every-ai-bot-hit-on-our-site#breadcrumb"},"dateModified":"2026-10-07T05:20:02.860Z","mainEntity":{"@id":"https://groas.com/post/we-counted-every-ai-bot-hit-on-our-site#article"}},{"@type":"BreadcrumbList","@id":"https://groas.com/post/we-counted-every-ai-bot-hit-on-our-site#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://groas.com/"},{"@type":"ListItem","position":2,"name":"Blog","item":"https://groas.com/blog"},{"@type":"ListItem","position":3,"name":"AI For Google Ads","item":"https://groas.com/category/ai-for-google-ads"},{"@type":"ListItem","position":4,"name":"We Counted AI Bot Hits for 30 Days. The Useful Number Was Fetches Per Page","item":"https://groas.com/post/we-counted-every-ai-bot-hit-on-our-site"}]},{"@type":"BlogPosting","@id":"https://groas.com/post/we-counted-every-ai-bot-hit-on-our-site#article","url":"https://groas.com/post/we-counted-every-ai-bot-hit-on-our-site","mainEntityOfPage":{"@id":"https://groas.com/post/we-counted-every-ai-bot-hit-on-our-site#webpage"},"isPartOf":{"@id":"https://groas.com/blog#blog"},"headline":"We Counted AI Bot Hits for 30 Days. The Useful Number Was Fetches Per Page","description":"Of the 13,716 requests claiming to be AI crawlers that hit /robots.txt on groas.com in 30 days, only 3,224 were verified. The other 10,492 were noise as far as this measurement is concerned: requests that presented a crawler identity we …","datePublished":"2026-10-07T05:20:02.860Z","dateModified":"2026-10-07T05:20:02.860Z","image":{"@type":"ImageObject","url":"https://groas.com/media/blog/ded929a3f8f0993b8f15bd4bb371f97a69d446561de97fc98e695256dd9048f8.png"},"author":{"@id":"https://groas.com/author/david#person"},"publisher":{"@id":"https://graph.groas.com/entity/groas"},"articleSection":"AI For Google Ads","keywords":"requests, verified, fetches, page, crawler, live-retrieval, robots txt, bot","wordCount":1859,"inLanguage":"en-US"},{"@type":"Person","@id":"https://groas.com/author/david#person","name":"David","jobTitle":"Founder & CEO @ groas","image":"https://groas.com/media/blog/363be9ad5255816654dac065d01bf1c6063110b400bb38cc199570267232027c.jpg","worksFor":{"@id":"https://graph.groas.com/entity/groas"}},{"@type":"FAQPage","@id":"https://groas.com/post/we-counted-every-ai-bot-hit-on-our-site#faq","mainEntity":[{"@type":"Question","name":"How many claimed AI crawler hits actually pass identity verification?","acceptedAnswer":{"@type":"Answer","text":"Of 13,716 requests claiming to be AI crawlers that hit /robots.txt on groas.com in 30 days, only 3,224 were verified, or 23.5%. The other 10,492, or 76.5%, carried a crawler identity that could not be validated and were not credited to the AI operator named in their User-Agent."}},{"@type":"Question","name":"Why can't I trust the User-Agent string in my server logs to identify AI bots?","acceptedAnswer":{"@type":"Answer","text":"A User-Agent is self-reported text, so any script can send User-Agent: ChatGPT-User or ClaudeBot, and a dashboard grouping by that string will count a fake visit. Research cited in the article found up to 98% of requests claiming certain AI crawler identities in raw logs can be fake. A claimed identity only counts as verified when the client IP matches the vendor-published IP ranges."}},{"@type":"Question","name":"What is the difference between a background crawl and a live-retrieval fetch?","acceptedAnswer":{"@type":"Answer","text":"A verified background crawl is a validated training or search-index crawler fetching a URL for a background task. A verified live-retrieval fetch is a validated user-retrieval agent, such as ChatGPT-User or Claude-User, fetching a page during its live-retrieval workflow. A live fetch still does not prove the page was quoted, cited, or shown to a buyer; the log records the fetch, not the finished answer."}},{"@type":"Question","name":"Do category pages get fetched by AI agents when they get lots of crawls?","acceptedAnswer":{"@type":"Answer","text":"On groas.com, category archives received 3,890 verified index crawls and zero verified live-retrieval fetches, while the blog index drew 4,220 crawls and just two live fetches. Two specific dated posts drew 157 and 149 live-retrieval fetches each. Crawl volume is therefore a poor substitute for the per-page live fetch count when judging AI visibility."}},{"@type":"Question","name":"Should robots.txt requests count as AI visits to my content?","acceptedAnswer":{"@type":"Answer","text":"No. The logs recorded 1,194 verified live-retrieval-agent requests to /robots.txt, and a rules-file fetch is not an article fetch. Counting both under a single live AI visits total makes content look busier than it was, so keep /robots.txt in its own row in your log analysis."}},{"@type":"Question","name":"What made the two posts attract live-retrieval fetches while the archives did not?","acceptedAnswer":{"@type":"Answer","text":"Both posts were dated and specific: they carried publication or modification timestamps, version details, and concrete configuration information, while an archive offers links and short descriptions. This is a mechanism, not a controlled experiment; the logs cannot isolate whether timestamps, page structure, or query subject caused the difference."}},{"@type":"Question","name":"How can I check which of my pages receive verified live-retrieval fetches?","acceptedAnswer":{"@type":"Answer","text":"Use a 30-day edge or server-log window with client IPs, paths, timestamps, and full User-Agent strings, and keep /robots.txt separate. Filter for claimed AI crawler identities, verify each client IP against vendor-published ranges such as openai.com/chatgpt-user.json, then compare verified index crawls and verified live-retrieval fetches by exact URL path."}}]}]}
```
