---
title: "We Counted 13,774 AI Bot Hits. Only 3,482 Verified."
description: "Our 30-day edge log shows why claimed AI bot traffic is a poor visibility metric: three category pages drew about 9,200 hits and no observed live-answer fetches, while two specific posts drew 380."
url: "https://groas.com/post/we-logged-30-days-of-ai-bots-hitting-our-2"
image: "https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/4568a1bc-01f3-4a1c-98db-2faea5fa1601.png"
published: "2026-10-11T05:26:08.843Z"
modified: "2026-10-11T05:26:08.909Z"
---

October 11, 2026 · 10 min read

# We Counted 13,774 AI Bot Hits. Only 3,482 Verified.

[Alexander PerelmanHead Of Product @ groas](https://groas.com/author/alexander-perelman)[LinkedIn](https://www.linkedin.com/in/alexander-433793253/)

![A doorwoman shines a UV torch on a queue of identically masked visitors, revealing that most are impostors and only a few are verified guests.](https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/4568a1bc-01f3-4a1c-98db-2faea5fa1601.png)

In this article

1. [Three counts that should never share a chart](#three-counts-that-should-never-share-a-chart)
2. [The 30-day log: big hit counts, small reading lists](#the-30-day-log-big-hit-counts-small-reading-lists)
3. [The pages assistants fetched instead](#the-pages-assistants-fetched-instead)
4. [How unverified and training traffic inflate the story](#how-unverified-and-training-traffic-inflate-the-story)
5. [What the fetched posts have in common](#what-the-fetched-posts-have-in-common)
6. [Measure the answers and the fetches](#measure-the-answers-and-the-fetches)
7. [The report I would replace this week](#the-report-i-would-replace-this-week)

Of the **13,774 hits to our robots.txt file that claimed to be AI bots, only 3,482 verified**. That is a 25.3% verification rate over 30 days of edge logs on our own site. If I had put the first number in an AI visibility report, roughly three-quarters of that chart would have represented unverified requests.

The page comparison is worse for the usual dashboard story. Three blog category pages drew about 9,200 claimed bot hits and **zero observed live-answer fetches**. Two dated, specific posts had a smaller crawl footprint but drew **380 verified live ChatGPT-User and Claude-User fetches** between them. Counting who knocked on the door tells you very little about what an assistant pulled while answering a person.

I used to tell clients to watch crawl volume as a proxy for AI visibility. I was wrong. The useful question is narrower: **Which verified assistant fetches reached a page during a live answer?** Even that does not prove the page appeared as a citation. It is a better consumption signal than a user-agent count, not a substitute for checking the answers themselves.

## Three counts that should never share a chart

A **claimed bot hit** is a request bearing a bot name in its user-agent. Any client can send `curl -A GPTBot` and appear in a log as GPTBot. [Verification means checking the client IP against a vendor-published range or confirming reverse-DNS in both directions](https://dev.to/inxprncd/how-to-verify-an-ai-crawler-is-who-it-says-it-is-470e). Across the 41 crawlers tracked in that reference, 20 document an IP-range check, 11 document reverse-DNS, and 16 publish no network method at all. A familiar name in a request header is not identity.

A **verified hit** passed an identity check. A **live assistant fetch** is narrower: a verified request from a fetcher such as ChatGPT-User or Claude-User retrieving material during an answer workflow, rather than a training crawler collecting pages for later. There is a further distinction the logs cannot erase. A fetch to `robots.txt` checks permissions; it is not a citation. A fetch to an article shows that the assistant retrieved it; it does not, by itself, show the exact text a person saw.

This is not a theoretical quibble. One cited measurement found that [16.7% of requests claiming to be ChatGPT-User were spoofed](https://github.com/anthropicbots/e-commerce/issues/494). Count the names without checking the clients and you count the costumes. **Verify identity first; separate permission checks from page fetches second.**

## The 30-day log: big hit counts, small reading lists

These figures come from our own edge log: 30 of 30 days measured, with bot identity checked by IP against vendor ranges. The path list available for this analysis was truncated, so the table is a comparison of the listed paths, not a full-site total. “Live” describes verified assistant-fetcher activity in the log, not independently confirmed citations in answer text.

| Path                                                |            Claimed bot hits |        Verified hits | Observed live assistant activity                       |
| --------------------------------------------------- | --------------------------: | -------------------: | ------------------------------------------------------ |
| `/robots.txt`                                       |                      13,774 |                3,482 | 1,195 Claude-User permission fetches                   |
| Homepage `/`                                        | Not the useful measure here | Verified subset only | 906 ChatGPT-User page fetches                          |
| Three blog category pages, combined                 |                 About 9,200 |                  687 | 0 observed live-answer fetches                         |
| Two dated posts: 2026 updates and AI Max data guide |       Small crawl footprint |             Verified | 380 ChatGPT-User and Claude-User page fetches combined |

Start with `robots.txt`, the easiest row to mistake for success. Its 13,774 claimed hits became 3,482 verified hits, a 25.3% verification rate. Within the verified activity were 1,195 live Claude-User fetches checking permissions. That is useful operational information: the fetcher reached the site and asked what it could access. It is not evidence that `robots.txt` supplied an answer to a buyer.

The unverified remainder sits in a messy web of scanners and impersonators. We saw probes for `/.env`, `/proc/self`, PHP-CGI paths and Kubernetes tokens wearing bot names. Other log analysis found that assistant fetchers claiming names such as ChatGPT-User and Claude-User [returned 404s at 43.7%, versus 10% for AI search crawlers](https://dev.to/lovedbyai/chatgpt-user-in-your-logs-how-to-tell-real-chatgpt-traffic-from-a-scanner-using-its-name-2jbp); 69% of the claimed ChatGPT-User 404s matched strict exploit paths. A raw hit chart can turn someone looking for an exposed file into someone supposedly interested in your content. That is quite a promotion.

Next come the three category pages. Together they drew about 9,200 claimed hits. Only 687 verified, a 7.5% verification rate, and none registered an observed live-answer fetch in this 30-day slice. A category page offers excerpts and links; the underlying post offers a complete passage. That is a plausible explanation for the split, not proof that every assistant always skips category pages. The outcome in this log is clear enough: **the pages with thousands of claimed hits supplied no observed live-answer page fetches.**

The broader crawl numbers show why volume needs careful handling. One reference reports [crawl-to-refer ratios of 23,951:1 for ClaudeBot, 1,276:1 for GPTBot and 111:1 for PerplexityBot, against 4.9:1 for Google](https://www.digitalapplied.com/blog/ai-crawler-bot-traffic-statistics-2026-data-reference) in early 2026. Those are not ratios from our site, and a referral is not the same as a citation. They do reinforce the narrower point: a crawl count and an audience outcome are different measures. **Do not rank pages for AI visibility by claimed bot hits.**

## The pages assistants fetched instead

The homepage drew **906 live ChatGPT-User page fetches** in the same 30 days. When someone asks about a business by name, a homepage that states what it does, whom it serves and how it charges gives a fetcher a direct place to look. Our log shows the fetches, not the prompts behind each one, so I would not assign every one of those 906 to a brand question. I would make sure the page answers that question plainly.

The two dated posts tell a different part of the story. A 2026 updates post and an AI Max data guide drew **380 verified live page fetches combined** across ChatGPT-User and Claude-User. Both lead with a self-contained answer, keep each paragraph focused, and include dates and numbers a reader can use. Compared with the category pages, they give an assistant a complete passage to retrieve rather than a menu of possible passages.

There is supporting context, with limits. A cited analysis reports that [ChatGPT-cited URLs were about 400 days newer than organic results for the same query, and pages updated within the previous 12 months were about twice as likely to be cited](https://www.linkedin.com/pulse/how-get-cited-ai-imvasa-rfndc). That does not establish why our two posts were fetched or promise that adding a date will earn a citation. It does make the dated, specific page a more sensible candidate for work than another broad archive. **Make the page useful to quote; do not just make the archive easy to crawl.**

## How unverified and training traffic inflate the story

User-agent spoofing is the first inflation source. A scanner can send a famous bot name and hope a dashboard trusts it. Behind Cloudflare or another proxy, the IP check can fail in either direction if you inspect the proxy address instead of the real client IP recorded in `CF-Connecting-IP`. [Check ChatGPT-User against OpenAI’s published `chatgpt-user.json` range, not the user-agent string](https://dev.to/lovedbyai/chatgpt-user-in-your-logs-how-to-tell-real-chatgpt-traffic-from-a-scanner-using-its-name-2jbp).

A person clicking through from a ChatGPT answer is a separate event. That visit arrives in a normal browser, potentially with a `chatgpt.com` referrer or `utm_source=chatgpt.com`, not with the ChatGPT-User bot name. A report built only from bot user-agents can include impostors while omitting the humans who actually reached the site. **Keep assistant fetches and referred human visits in separate columns.**

![Edge log showing claimed AI bot hits beside the smaller verified subset](https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/ab768bf9-6e90-4f0e-b3a4-e7ab3620752e.png)

Training crawls are the second inflation source. In the cited May 2026 purpose split, [51.8% of activity was training, 35.7% mixed and 9.3% search-only](https://www.digitalapplied.com/blog/ai-crawler-bot-traffic-statistics-2026-data-reference). Those figures are from that reference, not our edge log. A training crawl may matter for a different question, but it does not establish that your page was fetched for an answer this week. Add training volume to spoofed names and the raw bot-hit line can rise while your evidence of live use stays flat.

## What the fetched posts have in common

The useful distinction in our log is **complete answer versus index of answers**. The two posts that drew 380 live page fetches put an answer near the top of each section, then explain it. They keep one idea per paragraph and cover follow-up questions within the same page. The category pages introduce several posts but finish none of their arguments.

That structure fits the [passage-level citation pattern described in the cited analysis](https://www.linkedin.com/pulse/how-get-cited-ai-imvasa-rfndc): compact, self-contained passages are easier to use than fragments that send the reader elsewhere. The same analysis describes 3 to 10 citation slots per answer and about 14% overlap among top sources across ChatGPT, Perplexity and AI Overviews. Those figures are context, not an estimate of our citation share. Our logs tell us what was fetched; they cannot tell us whether any particular passage won one of those slots.

I would use the observed difference to set an editing order. Keep the homepage’s explanation of the business clear. Give dated, specific posts a direct answer before the supporting detail. Leave category pages to do their navigation job instead of trying to make their crawl counts look like demand. **Write the page an assistant can finish reading, not just the page a crawler can find.**

## Measure the answers and the fetches

I would put two views beside each other. The first checks what people might see: [run 20 to 30 buyer questions across ChatGPT, Perplexity, Gemini and Google AI Overviews](https://www.airops.com/blog/how-to-measure-ai-search-visibility), recording whether the brand appears, which pages are cited and which competitors appear alongside it. Repeat the questions on a cadence. The cited guidance notes that only about 30% of brands stay visible across consecutive runs, so one run is a thin basis for a trend.

The second view checks what reached the site: log the real client IP, verify the fetcher against the vendor’s published method, and report verified live page fetches by path. Keep permission requests such as `robots.txt`, training crawls, unverified names and human referrals visible, but separate. For the practical checks behind this view, [our curl and CDN audit covers robots.txt inspection, fetch tests and log verification](https://groas.com/post/can-chatgpt-actually-read-your-site-a-45).

The two views answer different questions. Prompt tracking shows whether an answer mentioned or cited you. Edge logs show whether a verified fetcher retrieved one of your pages. A mentioned brand with no matching page fetch in your log is a reason to investigate, not automatic proof of a content gap. A frequently fetched page that never appears in the answers you track deserves an editorial look, not an automatic rewrite. **Neither view should pretend to measure the other.**

## The report I would replace this week

If I inherited a site tomorrow, I would stop leading with total bot hits. I would make **verified live page fetches, by page** the main edge-log report and put the answers observed in prompt tracking next to it. The raw requests would remain available for security and troubleshooting. They would no longer pass as a visibility score.

Then I would work through the pages in this order:

1. **Check identity and purpose.** Record the real client IP, verify the vendor range where one is available, and separate live fetchers from training crawlers and permission checks before charting anything.
2. **Make the homepage answer the brand question.** State what the business sells, whom it serves and how pricing works in plain language. Keep that paragraph current rather than burying it in rotating hero copy.
3. **Give commercial pages an answer up front.** I would start with the five pages closest to money and put a complete, quotable answer before the explanation in each section. Use a date and a number where they genuinely help explain what has changed; a date pasted onto an unchanged page solves nothing.
4. **Stop judging category pages by bot traffic.** In our 30-day slice, they drew about 9,200 claimed hits and no observed live-answer fetches. They can still serve navigation without winning this particular count.

This is one site and 30 days, with a truncated path list. The log gives us a useful view of ChatGPT-User and Claude-User activity, but it cannot tie every fetch to a displayed citation. It says little directly about Gemini or Google AI Overviews, where an edge request cannot be assigned to a particular answer from these logs alone. That is why the prompt checks stay in the report.

The decision the numbers support is narrow. **Stop treating claimed bot hits as AI visibility. Verify the fetcher, distinguish page retrieval from permission and training traffic, then check the answers people actually receive.** In our log, that change moved attention away from the noisiest category pages and toward two specific posts that assistants fetched while answering. That is where I would spend the next editing hour.

## Frequently Asked Questions

### How many of the AI bot hits in the 30-day edge log were actually verified?

Only 3,482 of the 13,774 hits claiming to be AI bots verified successfully, a 25.3% verification rate over 30 days. That means roughly three-quarters of a raw bot-hit chart would represent unverified requests such as scanners and impersonators.

### How do I verify that an AI crawler in my logs is who it says it is?

Verify the client IP against a vendor-published range or confirm reverse-DNS in both directions. A familiar bot name in the user-agent string is not identity, since any client can send a header like GPTBot. For ChatGPT-User, check against OpenAI's published chatgpt-user.json range.

### Is a high AI bot hit count on a page a sign of AI visibility?

No. In the 30-day log, three blog category pages drew about 9,200 claimed bot hits with only 687 verified and zero observed live-answer fetches. Counting who knocked on the door says little about what an assistant actually pulled while answering a person.

### Which kinds of pages did AI assistants actually fetch during live answers?

Two dated, specific posts, a 2026 updates post and an AI Max data guide, drew 380 verified live ChatGPT-User and Claude-User page fetches combined, and the homepage drew 906 live ChatGPT-User page fetches. These pages put a self-contained answer near the top, while category pages only offer excerpts and links.

### Do robots.txt fetches from AI bots count as AI visibility?

No. A fetch to robots.txt checks what the fetcher is allowed to access, so it is not a citation and does not show that the file supplied an answer to a buyer. In the log, robots.txt drew 13,774 claimed hits, 3,482 verified, including 1,195 live Claude-User permission fetches.

### Does a human clicking through from a ChatGPT answer show up as ChatGPT-User traffic in logs?

No. A person clicking through arrives in a normal browser, possibly with a chatgpt.com referrer or utm_source=chatgpt.com, not with the ChatGPT-User bot name. A report built only from bot user-agents can include impostors while omitting the humans who actually reached the site, so keep the two in separate columns.

### How should I measure AI visibility instead of counting bot hits?

Use two views side by side. Run 20 to 30 buyer questions across ChatGPT, Perplexity, Gemini and Google AI Overviews on a regular cadence, recording brand mentions, cited pages and competitors. In parallel, log the real client IP, verify fetchers against vendor-published methods, and report verified live page fetches by path, keeping permission requests, training crawls, unverified names and human referrals separate.

## Pay For Results, Not For Hours

Businesses buy the outcome, agencies resell it, and groas answers for it either way.

[See If You Qualify](https://groas.typeform.com/to/xC1bQNUT)

## Structured data

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://graph.groas.com/entity/groas","name":"groas","url":"https://groas.com/","logo":{"@type":"ImageObject","url":"https://groas.com/icon-512.png","width":512,"height":512},"sameAs":["https://www.linkedin.com/company/groas/"]},{"@type":"WebSite","@id":"https://groas.com/#website","name":"groas","url":"https://groas.com/","inLanguage":"en-US","publisher":{"@id":"https://graph.groas.com/entity/groas"}},{"@type":"WebPage","@id":"https://groas.com/post/we-logged-30-days-of-ai-bots-hitting-our-2#webpage","url":"https://groas.com/post/we-logged-30-days-of-ai-bots-hitting-our-2","name":"We Counted 13,774 AI Bot Hits. Only 3,482 Verified.","description":"Our 30-day edge log shows why claimed AI bot traffic is a poor visibility metric: three category pages drew about 9,200 hits and no observed live-answer fetches, while two specific posts drew 380.","inLanguage":"en-US","isPartOf":{"@id":"https://groas.com/#website"},"about":{"@id":"https://graph.groas.com/entity/groas"},"primaryImageOfPage":{"@type":"ImageObject","url":"https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/4568a1bc-01f3-4a1c-98db-2faea5fa1601.png"},"breadcrumb":{"@id":"https://groas.com/post/we-logged-30-days-of-ai-bots-hitting-our-2#breadcrumb"},"dateModified":"2026-10-11T05:26:08.909Z","mainEntity":{"@id":"https://groas.com/post/we-logged-30-days-of-ai-bots-hitting-our-2#article"}},{"@type":"BreadcrumbList","@id":"https://groas.com/post/we-logged-30-days-of-ai-bots-hitting-our-2#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://groas.com/"},{"@type":"ListItem","position":2,"name":"Blog","item":"https://groas.com/blog"},{"@type":"ListItem","position":3,"name":"We Counted 13,774 AI Bot Hits. Only 3,482 Verified.","item":"https://groas.com/post/we-logged-30-days-of-ai-bots-hitting-our-2"}]},{"@type":"BlogPosting","@id":"https://groas.com/post/we-logged-30-days-of-ai-bots-hitting-our-2#article","url":"https://groas.com/post/we-logged-30-days-of-ai-bots-hitting-our-2","mainEntityOfPage":{"@id":"https://groas.com/post/we-logged-30-days-of-ai-bots-hitting-our-2#webpage"},"isPartOf":{"@id":"https://groas.com/blog#blog"},"headline":"We Counted 13,774 AI Bot Hits. Only 3,482 Verified.","description":"Our 30-day edge log shows why claimed AI bot traffic is a poor visibility metric: three category pages drew about 9,200 hits and no observed live-answer fetches, while two specific posts drew 380.","datePublished":"2026-10-11T05:26:08.843Z","dateModified":"2026-10-11T05:26:08.909Z","image":{"@type":"ImageObject","url":"https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/4568a1bc-01f3-4a1c-98db-2faea5fa1601.png"},"author":{"@id":"https://groas.com/author/alexander-perelman#person"},"publisher":{"@id":"https://graph.groas.com/entity/groas"},"keywords":"verified, page, bot, fetches, hits, pages, log, bot hits","wordCount":2099,"inLanguage":"en-US"},{"@type":"Person","@id":"https://groas.com/author/alexander-perelman#person","name":"Alexander Perelman","jobTitle":"Head Of Product @ groas","description":"Ex Goldman Sachs and Ex Stanford Computer Science","image":"https://groas.com/media/blog/4f17cac3a81acc224cc1b2cfab8f94bbc295a3089d13e8421125771ef8d4064f.jpg","worksFor":{"@id":"https://graph.groas.com/entity/groas"},"sameAs":["https://www.linkedin.com/in/alexander-433793253/"]},{"@type":"FAQPage","@id":"https://groas.com/post/we-logged-30-days-of-ai-bots-hitting-our-2#faq","mainEntity":[{"@type":"Question","name":"How many of the AI bot hits in the 30-day edge log were actually verified?","acceptedAnswer":{"@type":"Answer","text":"Only 3,482 of the 13,774 hits claiming to be AI bots verified successfully, a 25.3% verification rate over 30 days. That means roughly three-quarters of a raw bot-hit chart would represent unverified requests such as scanners and impersonators."}},{"@type":"Question","name":"How do I verify that an AI crawler in my logs is who it says it is?","acceptedAnswer":{"@type":"Answer","text":"Verify the client IP against a vendor-published range or confirm reverse-DNS in both directions. A familiar bot name in the user-agent string is not identity, since any client can send a header like GPTBot. For ChatGPT-User, check against OpenAI's published chatgpt-user.json range."}},{"@type":"Question","name":"Is a high AI bot hit count on a page a sign of AI visibility?","acceptedAnswer":{"@type":"Answer","text":"No. In the 30-day log, three blog category pages drew about 9,200 claimed bot hits with only 687 verified and zero observed live-answer fetches. Counting who knocked on the door says little about what an assistant actually pulled while answering a person."}},{"@type":"Question","name":"Which kinds of pages did AI assistants actually fetch during live answers?","acceptedAnswer":{"@type":"Answer","text":"Two dated, specific posts, a 2026 updates post and an AI Max data guide, drew 380 verified live ChatGPT-User and Claude-User page fetches combined, and the homepage drew 906 live ChatGPT-User page fetches. These pages put a self-contained answer near the top, while category pages only offer excerpts and links."}},{"@type":"Question","name":"Do robots.txt fetches from AI bots count as AI visibility?","acceptedAnswer":{"@type":"Answer","text":"No. A fetch to robots.txt checks what the fetcher is allowed to access, so it is not a citation and does not show that the file supplied an answer to a buyer. In the log, robots.txt drew 13,774 claimed hits, 3,482 verified, including 1,195 live Claude-User permission fetches."}},{"@type":"Question","name":"Does a human clicking through from a ChatGPT answer show up as ChatGPT-User traffic in logs?","acceptedAnswer":{"@type":"Answer","text":"No. A person clicking through arrives in a normal browser, possibly with a chatgpt.com referrer or utm_source=chatgpt.com, not with the ChatGPT-User bot name. A report built only from bot user-agents can include impostors while omitting the humans who actually reached the site, so keep the two in separate columns."}},{"@type":"Question","name":"How should I measure AI visibility instead of counting bot hits?","acceptedAnswer":{"@type":"Answer","text":"Use two views side by side. Run 20 to 30 buyer questions across ChatGPT, Perplexity, Gemini and Google AI Overviews on a regular cadence, recording brand mentions, cited pages and competitors. In parallel, log the real client IP, verify fetchers against vendor-published methods, and report verified live page fetches by path, keeping permission requests, training crawls, unverified names and human referrals separate."}}]}]}
```
