Our homepage logged 11,441 hits claiming to be AI bots in 30 days. Only 2,187 were verified. I used to tell clients to watch bot hits in the logs as a proxy for AI visibility. I was wrong.

The rest of the month made the distinction harder to ignore. Category pages drew more than 3,000 claimed hits each and fed zero live answers. One specific, dated YouTube ads guide had 294 verified fetches; 273 fed live ChatGPT and Claude answers. AI attention and AI citation are different things. Our logs point to a page-shape problem, not a crawl-volume problem.

What the edge logs count, and what they cannot see

This is a 30-day slice of groas.com edge logs, not Google Search Console and not a rank tracker. I split the activity into three groups:

  1. Claimed AI hits: requests carrying an AI bot user-agent.
  2. Verified hits: requests whose IP matched the relevant vendor’s published list.
  3. Live-answer fetches: user-triggered retrieval requests associated with a live ChatGPT, Claude, or Perplexity answer.

Those are different denominators. Any script can claim to be ChatGPT-User in its user-agent. IP verification gets us closer to genuine vendor traffic. It does not make every verified request an answer, and an answer-related fetch does not, by itself, prove that a reader saw a visible citation. The logs tell us which pages fed live answers. They cannot show the full text of every answer or settle why a model chose one page over another.

The distinction matters before we even compare pages. Training crawlers gather material for models; retrieval agents fetch material in response to a question. Counting training and retrieval bots together turns two different jobs into one impressive-looking traffic line. Impressive-looking is not a reporting category I trust.

Verification also needs to be maintained, not performed once and forgotten. A separate two-week sample across 16 AI crawlers found that 5.7% of traffic claiming an AI crawler identity was spoofed; about one in six requests claiming to be ChatGPT-User was fake. That is not the spoofing rate in our logs. It is a useful reason to check identity rather than trust a label. For OpenAI, that means matching IPs against its per-agent lists on a refresh schedule. For Anthropic, the published list has changed from its older tokens. A report built on user-agents alone counts strangers knocking on the door as invited guests.

Finding 1: the homepage’s 11,441 claimed hits became 2,187 verified hits

Here is the cut that changed how I read a bot report. Of 11,441 homepage hits claiming an AI identity, 2,187 survived IP verification: 19.1%. Category pages had a lower verified share, roughly 7–8%. Their 3,000-plus figures are claimed hits, not verified fetches. That distinction is easy to lose when both appear on a chart labelled “AI traffic.”

Page groupClaimed AI hitsVerified hitsVerified share
Homepage11,4412,18719.1%
Category pages3,000+ eachNot reported as a count~7–8%

Every figure in this table comes from the same 30-day groas edge-log slice. The category-page range is approximate; the dataset here does not give a per-page verified count to put in that column.

Say you spend $20k a month and get a report celebrating a climb in crawler hits. Before anyone calls it visibility, ask whether those hits were checked against vendor IPs. On our homepage, roughly four out of five claimed hits did not pass that check. I would not apply that exact fraction to every site, or even to every page on ours. I would stop treating an unchecked user-agent chart as evidence that an AI service is paying attention.

First verify the visitor. Then decide whether the visit means anything.

Finding 2: category pages got thousands of hits and fed zero answers

The category pages each drew more than 3,000 claimed bot hits during the window. None fed a live answer in this dataset. Same site, same month, very different outcome from the dated guides. A hit says something requested a URL. It does not tell you whether the request came from a genuine AI agent, whether that agent was gathering training material, or whether a user had asked a question.

The bot’s job matters here. OpenAI uses separate robots.txt tokens for GPTBot, OAI-SearchBot, and ChatGPT-User: training, search indexing, and a page pull related to a user request. Blocking GPTBot is not the same as blocking OAI-SearchBot. Perplexity likewise distinguishes PerplexityBot, which indexes, from Perplexity-User, which fetches for a user. One total labelled “AI crawls” hides decisions that affect different parts of the path to an answer.

Bot typeJobRelationship to an answer
GPTBot, ClaudeBotTraining crawlNot a live-answer fetch
OAI-SearchBot, PerplexityBotSearch indexingHelps make pages retrievable
ChatGPT-User, Claude-User, Perplexity-UserUser-triggered retrievalCan feed a live answer; not proof of a visible citation

The table describes bot roles, not counts from our logs. It also explains why I would not respond to the category-page result by chasing more raw crawls. More requests of the wrong kind would make the top line bigger while leaving the answer line at zero.

Report training, indexing, and live retrieval separately. Otherwise, the number most likely to impress someone in a meeting is the least useful one for this decision.

Finding 3: one dated guide fed 273 live answers

The YouTube ads 2026 guide had 294 verified fetches, 273 of which fed a live answer. That is 92.9% of its verified fetches in this 30-day slice. The Google Ads updates guide fed 106 live answers, and the AI Max guide fed 139. We do not have verified-fetch totals for those last two guides in this cut, so I cannot give them comparable rates.

PageVerified fetchesLive-answer fetchesShare of verified fetches feeding answers
YouTube ads 2026 guide29427392.9%
AI Max guideNot available139Not available
Google Ads updates guideNot available106Not available
Category pagesNot available0Not available

These are groas edge-log counts. The 3,000-plus category-page hits in the previous table were claimed hits, so dividing zero answers by them would not produce the same rate shown for the YouTube guide. Zero is still zero. The missing denominators are still missing.

What did the pages feeding answers have that the category pages did not? The guides name a specific subject, carry a visible date, and put numbers, steps, and definitions in the body. A category page lists routes to other pages. If a retrieval agent needs material for a narrow question, the guide offers an answer it can use; the index asks it to keep looking. That is a plausible explanation for the gap, not a controlled test proving that a year in a title caused 273 answers.

Freshness deserves attention, with the same caution. One analysis of generative-AI citations reports that 95% of cited pages had been updated within 10 months and pages with a visible last-updated date earned 1.8 times more citations. Another synthesis reports 79% day-to-day churn in ChatGPT sources and says more than 70% of cited pages had been updated within the previous 12 months. Those are outside analyses, not measurements from our edge logs, and neither turns a date stamp into a guaranteed citation.

I used to treat the date stamp as cosmetic. Now I would check whether the date reflects a real update, then refresh the numbers and steps that changed. Rewriting an introduction while leaving stale facts in place is tidying the shop window and ignoring the shelves. Make the answer specific and current; do not merely make the page look current.

Finding 4: permission checks are traffic, not citations

Our robots.txt received 1,160 Claude-User requests classified as live-answer fetches in the window. Nobody needs a quote from a robots.txt file in a ChatGPT answer. That count is a reminder that a retrieval path can generate requests to supporting files as well as to pages a reader might actually see.

I would not count those 1,160 requests as 1,160 citations, or use them to calculate a page’s answer rate. I would use them to check a less glamorous part of the setup: is the file available, current, and explicit about the agents we intend to allow? A stale block can interfere with the path we want open. Guessing from an old list of bot names is not a substitute for checking current vendor documentation and verified logs.

The homepage needs the same restraint. It receives plenty of checks, and its verified share sits near one in five, but the logs do not show it performing like the guides as a source for live answers. Scale makes the traffic feel persuasive: an industry account reports roughly 50 billion AI crawler requests a day in March 2025, while Googlebot still reached more URLs than ClaudeBot or PerplexityBot. Those are ecosystem figures, not a multiplier for our site. They help explain why a busy server log and a thin stream of useful retrieval can coexist.

Keep access working, but do not mistake access checks for demand for the homepage.

The decision: fix the measurement, then fix one page

If I were looking at this on a business site, I would change the report before changing the content. The order matters:

  1. Verify claimed bots by IP. Keep the vendor lists current. A user-agent-only trend is not a reliable starting point.
  2. Split requests by job. Separate training crawls, search indexing, and user-triggered retrieval instead of presenting one crawler total.
  3. Compare pages using matching denominators. For pages with both figures, compare verified fetches with fetches that fed live answers. Mark missing counts as missing rather than filling the gap with a claimed-hit total.
  4. Choose one narrow guide to improve. Give it a question it can answer directly, a date that reflects real maintenance, and numbers, definitions, or steps that do the work on the page. Then watch the same measures over time.

Crawl access is part of that work, but the controls are not interchangeable. If the aim is to keep OpenAI’s training crawler out while leaving its search and user-request paths open, the documented distinction among GPTBot, OAI-SearchBot, and ChatGPT-User matters. Check the rules for each service rather than copying an OpenAI rule onto Perplexity or Claude. A broad block may be simpler to type and harder to diagnose later.

The content decision is narrower than “publish more AI-friendly pages.” Our YouTube guide did not stand out because the logs show it was long, and the logs cannot tell us that a date alone made it useful. It stood out because verified agents fetched it and it repeatedly fed live answers, while category pages with thousands of claimed hits fed none. I would use that contrast to inspect the guide’s structure and update another page that can answer a similarly specific question. I would not turn a category index into a longer category index and call it a test.

The limits: one site, top paths, 30 days

This is one site with an ads-focused audience, observed for 30 days. The path list available for this analysis was truncated to top paths. One-off fetches in the long tail are missing, so I cannot say these are the only guides that fed answers. I also cannot turn a correlation between page shape and live-answer fetches into proof that formatting alone caused the difference.

The window is long enough to expose a stark contrast in this site’s logs, but the exact counts can move. The outside tracking synthesis reports day-to-day source churn; a different month could change which guides appear in the table. Treat our rates as a prompt to examine your own pages, not benchmarks to paste into a forecast.

There is a more basic limit, too: a fetch that fed an answer is not a record of a visible citation. The logs support a decision about which pages and requests to investigate. They do not let me promise that every answer displayed a link, or that making another page look like the YouTube guide will reproduce its rate.

The decision they do support is clear. Verify bot traffic before reporting it. Separate live retrieval from the other work bots do. Then put the next content effort into a page that can answer a specific question with current, extractable details. If a page has no fresh numbers, definitions, or steps to offer, I would skip the cosmetic rewrite. It may attract another hit without giving an answer anything useful to carry.

Put the month in one picture and it looks like a funnel, not a flood. Thousands of claimed hits arrive at the top. IP checks remove most homepage claims. Bot roles separate background activity from user-triggered retrieval. At the bottom, a dated guide feeds answers while category pages with heavy claimed traffic feed none. That is why I stopped reporting crawls as visibility.

Funnel from claimed AI bot hits through IP verification to live-answer fetches

The pages that made it through looked nothing like a homepage or a category index. They looked like work: a year in the title, a date on the page, numbers and steps an answer could use without guessing. Build that shape once, then let verified logs tell you which question the page actually answers.

Dated 2026 guide beside a chat answer