Of the 13,774 hits to our robots.txt file that claimed to be AI bots, only 3,482 verified. That is a 25.3% verification rate over 30 days of edge logs on our own site. If I had put the first number in an AI visibility report, roughly three-quarters of that chart would have represented unverified requests.

The page comparison is worse for the usual dashboard story. Three blog category pages drew about 9,200 claimed bot hits and zero observed live-answer fetches. Two dated, specific posts had a smaller crawl footprint but drew 380 verified live ChatGPT-User and Claude-User fetches between them. Counting who knocked on the door tells you very little about what an assistant pulled while answering a person.

I used to tell clients to watch crawl volume as a proxy for AI visibility. I was wrong. The useful question is narrower: Which verified assistant fetches reached a page during a live answer? Even that does not prove the page appeared as a citation. It is a better consumption signal than a user-agent count, not a substitute for checking the answers themselves.

Three counts that should never share a chart

A claimed bot hit is a request bearing a bot name in its user-agent. Any client can send curl -A GPTBot and appear in a log as GPTBot. Verification means checking the client IP against a vendor-published range or confirming reverse-DNS in both directions. Across the 41 crawlers tracked in that reference, 20 document an IP-range check, 11 document reverse-DNS, and 16 publish no network method at all. A familiar name in a request header is not identity.

A verified hit passed an identity check. A live assistant fetch is narrower: a verified request from a fetcher such as ChatGPT-User or Claude-User retrieving material during an answer workflow, rather than a training crawler collecting pages for later. There is a further distinction the logs cannot erase. A fetch to robots.txt checks permissions; it is not a citation. A fetch to an article shows that the assistant retrieved it; it does not, by itself, show the exact text a person saw.

This is not a theoretical quibble. One cited measurement found that 16.7% of requests claiming to be ChatGPT-User were spoofed. Count the names without checking the clients and you count the costumes. Verify identity first; separate permission checks from page fetches second.

The 30-day log: big hit counts, small reading lists

These figures come from our own edge log: 30 of 30 days measured, with bot identity checked by IP against vendor ranges. The path list available for this analysis was truncated, so the table is a comparison of the listed paths, not a full-site total. “Live” describes verified assistant-fetcher activity in the log, not independently confirmed citations in answer text.

PathClaimed bot hitsVerified hitsObserved live assistant activity
/robots.txt13,7743,4821,195 Claude-User permission fetches
Homepage /Not the useful measure hereVerified subset only906 ChatGPT-User page fetches
Three blog category pages, combinedAbout 9,2006870 observed live-answer fetches
Two dated posts: 2026 updates and AI Max data guideSmall crawl footprintVerified380 ChatGPT-User and Claude-User page fetches combined

Start with robots.txt, the easiest row to mistake for success. Its 13,774 claimed hits became 3,482 verified hits, a 25.3% verification rate. Within the verified activity were 1,195 live Claude-User fetches checking permissions. That is useful operational information: the fetcher reached the site and asked what it could access. It is not evidence that robots.txt supplied an answer to a buyer.

The unverified remainder sits in a messy web of scanners and impersonators. We saw probes for /.env, /proc/self, PHP-CGI paths and Kubernetes tokens wearing bot names. Other log analysis found that assistant fetchers claiming names such as ChatGPT-User and Claude-User returned 404s at 43.7%, versus 10% for AI search crawlers; 69% of the claimed ChatGPT-User 404s matched strict exploit paths. A raw hit chart can turn someone looking for an exposed file into someone supposedly interested in your content. That is quite a promotion.

Next come the three category pages. Together they drew about 9,200 claimed hits. Only 687 verified, a 7.5% verification rate, and none registered an observed live-answer fetch in this 30-day slice. A category page offers excerpts and links; the underlying post offers a complete passage. That is a plausible explanation for the split, not proof that every assistant always skips category pages. The outcome in this log is clear enough: the pages with thousands of claimed hits supplied no observed live-answer page fetches.

The broader crawl numbers show why volume needs careful handling. One reference reports crawl-to-refer ratios of 23,951:1 for ClaudeBot, 1,276:1 for GPTBot and 111:1 for PerplexityBot, against 4.9:1 for Google in early 2026. Those are not ratios from our site, and a referral is not the same as a citation. They do reinforce the narrower point: a crawl count and an audience outcome are different measures. Do not rank pages for AI visibility by claimed bot hits.

The pages assistants fetched instead

The homepage drew 906 live ChatGPT-User page fetches in the same 30 days. When someone asks about a business by name, a homepage that states what it does, whom it serves and how it charges gives a fetcher a direct place to look. Our log shows the fetches, not the prompts behind each one, so I would not assign every one of those 906 to a brand question. I would make sure the page answers that question plainly.

The two dated posts tell a different part of the story. A 2026 updates post and an AI Max data guide drew 380 verified live page fetches combined across ChatGPT-User and Claude-User. Both lead with a self-contained answer, keep each paragraph focused, and include dates and numbers a reader can use. Compared with the category pages, they give an assistant a complete passage to retrieve rather than a menu of possible passages.

There is supporting context, with limits. A cited analysis reports that ChatGPT-cited URLs were about 400 days newer than organic results for the same query, and pages updated within the previous 12 months were about twice as likely to be cited. That does not establish why our two posts were fetched or promise that adding a date will earn a citation. It does make the dated, specific page a more sensible candidate for work than another broad archive. Make the page useful to quote; do not just make the archive easy to crawl.

How unverified and training traffic inflate the story

User-agent spoofing is the first inflation source. A scanner can send a famous bot name and hope a dashboard trusts it. Behind Cloudflare or another proxy, the IP check can fail in either direction if you inspect the proxy address instead of the real client IP recorded in CF-Connecting-IP. Check ChatGPT-User against OpenAI’s published chatgpt-user.json range, not the user-agent string.

A person clicking through from a ChatGPT answer is a separate event. That visit arrives in a normal browser, potentially with a chatgpt.com referrer or utm_source=chatgpt.com, not with the ChatGPT-User bot name. A report built only from bot user-agents can include impostors while omitting the humans who actually reached the site. Keep assistant fetches and referred human visits in separate columns.

Edge log showing claimed AI bot hits beside the smaller verified subset

Training crawls are the second inflation source. In the cited May 2026 purpose split, 51.8% of activity was training, 35.7% mixed and 9.3% search-only. Those figures are from that reference, not our edge log. A training crawl may matter for a different question, but it does not establish that your page was fetched for an answer this week. Add training volume to spoofed names and the raw bot-hit line can rise while your evidence of live use stays flat.

What the fetched posts have in common

The useful distinction in our log is complete answer versus index of answers. The two posts that drew 380 live page fetches put an answer near the top of each section, then explain it. They keep one idea per paragraph and cover follow-up questions within the same page. The category pages introduce several posts but finish none of their arguments.

That structure fits the passage-level citation pattern described in the cited analysis: compact, self-contained passages are easier to use than fragments that send the reader elsewhere. The same analysis describes 3 to 10 citation slots per answer and about 14% overlap among top sources across ChatGPT, Perplexity and AI Overviews. Those figures are context, not an estimate of our citation share. Our logs tell us what was fetched; they cannot tell us whether any particular passage won one of those slots.

I would use the observed difference to set an editing order. Keep the homepage’s explanation of the business clear. Give dated, specific posts a direct answer before the supporting detail. Leave category pages to do their navigation job instead of trying to make their crawl counts look like demand. Write the page an assistant can finish reading, not just the page a crawler can find.

Measure the answers and the fetches

I would put two views beside each other. The first checks what people might see: run 20 to 30 buyer questions across ChatGPT, Perplexity, Gemini and Google AI Overviews, recording whether the brand appears, which pages are cited and which competitors appear alongside it. Repeat the questions on a cadence. The cited guidance notes that only about 30% of brands stay visible across consecutive runs, so one run is a thin basis for a trend.

The second view checks what reached the site: log the real client IP, verify the fetcher against the vendor’s published method, and report verified live page fetches by path. Keep permission requests such as robots.txt, training crawls, unverified names and human referrals visible, but separate. For the practical checks behind this view, our curl and CDN audit covers robots.txt inspection, fetch tests and log verification.

The two views answer different questions. Prompt tracking shows whether an answer mentioned or cited you. Edge logs show whether a verified fetcher retrieved one of your pages. A mentioned brand with no matching page fetch in your log is a reason to investigate, not automatic proof of a content gap. A frequently fetched page that never appears in the answers you track deserves an editorial look, not an automatic rewrite. Neither view should pretend to measure the other.

The report I would replace this week

If I inherited a site tomorrow, I would stop leading with total bot hits. I would make verified live page fetches, by page the main edge-log report and put the answers observed in prompt tracking next to it. The raw requests would remain available for security and troubleshooting. They would no longer pass as a visibility score.

Then I would work through the pages in this order:

  1. Check identity and purpose. Record the real client IP, verify the vendor range where one is available, and separate live fetchers from training crawlers and permission checks before charting anything.
  2. Make the homepage answer the brand question. State what the business sells, whom it serves and how pricing works in plain language. Keep that paragraph current rather than burying it in rotating hero copy.
  3. Give commercial pages an answer up front. I would start with the five pages closest to money and put a complete, quotable answer before the explanation in each section. Use a date and a number where they genuinely help explain what has changed; a date pasted onto an unchanged page solves nothing.
  4. Stop judging category pages by bot traffic. In our 30-day slice, they drew about 9,200 claimed hits and no observed live-answer fetches. They can still serve navigation without winning this particular count.

This is one site and 30 days, with a truncated path list. The log gives us a useful view of ChatGPT-User and Claude-User activity, but it cannot tie every fetch to a displayed citation. It says little directly about Gemini or Google AI Overviews, where an edge request cannot be assigned to a particular answer from these logs alone. That is why the prompt checks stay in the report.

The decision the numbers support is narrow. Stop treating claimed bot hits as AI visibility. Verify the fetcher, distinguish page retrieval from permission and training traffic, then check the answers people actually receive. In our log, that change moved attention away from the noisiest category pages and toward two specific posts that assistants fetched while answering. That is where I would spend the next editing hour.