Our logs recorded 13,405 robots.txt requests from claimed AI user agents in 30 days. Only 3,176 came from verified search bots. If you are about to pay someone to fix your AI visibility, start with that gap. A bot name in a log does not mean an answer engine read your pages, and a longer technical checklist will not make an unreadable page readable. Too much of this market sells familiar web hygiene with an AI label. Here are the seven myths I would clear away before spending on another audit.
Myth 1: “Bot hits in your logs mean AI is reading your site”
A spike in GPTBot or ClaudeBot entries looks like progress. I understand the temptation to open a prompt-monitoring dashboard and wait for the citations to arrive. But a user-agent string is a claim made by whoever sent the request, not proof of identity. As Cloudflare explains in its discussion of AI crawler traffic and evasive scrapers, scrapers can present themselves as legitimate crawlers.
We saw the distinction when we logged AI bot hits on our own domain: of 13,405 robots.txt requests from claimed AI user agents, only 3,176 were verified search bots. That leaves 10,229 requests you should not count as evidence that an answer engine is interested in your content. And even a verified request for robots.txt tells you nothing, by itself, about whether the bot fetched a product page or could read its text.
Verify the visitor, then inspect the page request. A graph of bot names is not a citation pipeline. It may just be a graph of scrapers that know how to spell GPTBot.
Myth 2: “Every AI crawler renders your JavaScript”
Googlebot can render JavaScript, so teams assume every crawler arriving at a React site sees what a person sees. That assumption can turn an expensive page rebuild into an even more expensive distraction: the page looks finished in a browser, while its initial HTML contains little more than <div id="root"></div>.
An analysis of more than 500 million GPTBot requests found no evidence of client-side JavaScript execution. GPTBot fetched .js files in 11.5% of visits, and ClaudeBot fetched them in 23.8% of cases. Fetching a script is not the same as running it. If the text that explains your product arrives only after that script runs, you cannot assume those crawlers receive the explanation.

This is where the platforms diverge. As Yar’s analysis of AI bot rendering explains, Google AI Overviews and AI Mode can draw on Google’s rendered search index. Other answer engines use their own fetching paths. A page can therefore be discoverable through Google and still present an empty shell to a different crawler. The browser view alone cannot settle the question; it shows what happened after scripts ran. Check the initial server response for the words you need cited. If they are missing, tweaking the client-side layout will not put them there.
Myth 3: “JSON-LD schema is your ticket into AI citations”
Schema is easy to sell because it looks like an instruction manual for machines. An audit flags missing markup, an agency adds an elaborate JSON-LD array, and everyone can point to a completed task. The leap from machine-readable markup to guaranteed citation is where the pitch falls apart.
In a study of 1,885 pages that added JSON-LD against 4,000 matched controls, Ahrefs found a 4.6% relative drop in Google AI Overview citations. The measured changes for ChatGPT (+2.2%) and Google AI Mode (+2.4%) were statistically indistinguishable from zero. Those results do not say schema is useless for every search purpose. They do say you should not buy a semantic entity audit on the promise that markup will make an answer engine cite you.
The more basic question comes first: does the readable page actually answer what the user asked? If a comparison page never states the difference between two products in its body text, wrapping the page in immaculate structured data does not supply the missing answer. Treat schema as markup, not a substitute for substance.
Myth 4: “Adding an llms.txt file is the technical shortcut”
An llms.txt file is appealing for the same reason a 42-page audit is appealing: it offers a tidy artifact. Put a Markdown index at the root, list your best pages, and the bots will supposedly follow along. Generating the file is easier than fixing a page that sends an empty HTML shell, which may explain the enthusiasm.
The observed fetching is less exciting. In a 12-week study of 83 domains with llms.txt files, OpenAI crawlers requested robots.txt 3,990 times and llms.txt seven times. ClaudeBot requested them 3,120 and nine times, respectively; PerplexityBot requested robots.txt 775 times and llms.txt zero times. Another 90-day benchmark of 62,100 AI crawler requests put llms.txt requests at 0.13% of the total.
There is a narrower idea worth separating from the manifest pitch. Snorklee’s log benchmarks found requests for guessable Markdown page URLs such as /page.md: 34.8% for GPTBot and 22.7% for OAI-SearchBot in the reported comparisons. A clean text version of a valuable page may be useful where it is actually fetched. It still needs to contain the answer. Do not pay for a manifest as though deploying it were the same as getting your pages read. Start with the pages.
Myth 5: “Our landing pages convert paid traffic, so they are ready for AI search”
I have spent enough time with PPC landing pages to respect the tradeoff. A page built for a click often hides detail to keep a person moving: specifications sit in accordions, feature comparisons live in tabs, and pricing depends on an interactive calculator. That can be sensible for paid traffic. It can also leave a non-rendering crawler with a headline, a testimonial, and a form. None of those elements explains a buried specification or comparison on its own.
The issue is not that every collapsed accordion is invisible. It is how its contents get into the page. If the text is already in the initial HTML, a crawler can encounter it without clicking anything. If a script inserts the text only after a click, a fetcher that does not execute that interaction will miss it. The distinction matters more than whether the accordion looks open in your browser. The analysis of how bots process hidden accordion content is a useful prompt to inspect the HTML rather than judge the page by its design.
A paid landing page and an answerable page can share a goal without sharing every design choice. Keep the facts a crawler needs in the server response. The form can stay; the explanation cannot depend on someone opening Tab Four.
Myth 6: “A content-gap report tells us what to write for AI citations”
A content-gap spreadsheet feels decisive because it gives you a list. Too often, the list is familiar SEO phrase-matching work with AI added to the product name. It tells you to mention “enterprise scaling” or “turnkey integration” more often, then calls the result a visibility strategy.
If a user asks why an integration breaks, another repetition of “enterprise scaling” is not an answer. A passage that names the failure and explains the mechanism gives an answer engine something it can use. I would rather fix one page that buries that explanation beneath eight sections of setup than commission twenty pages built around suggested phrases. Find the unanswered question, not just the missing term.
Myth 7: “Blocking or allowing AI bots is a one-time robots.txt toggle”
This is the hardest myth to kill because a toggle feels like control. Someone wants to keep training crawlers away from proprietary material, so they block every AI-sounding user agent. Someone else wants citations, so they paste in an allow-all rule. Both decisions treat AI crawlers as one job with one consequence.
They are not. As this breakdown of AI crawler robots.txt configurations lays out, the crawler named in a rule matters:
| Function | Examples | What the rule affects |
|---|---|---|
| Model training | GPTBot, ClaudeBot; Google-Extended is a control token | Whether the relevant operator may use permitted content for training or related purposes. A robots.txt rule is not access control for proprietary data. |
| Search indexing | OAI-SearchBot, Claude-SearchBot, PerplexityBot | Whether those crawlers can fetch pages for their search experiences. Blocking them can reduce direct visibility there. |
| User-triggered fetches | ChatGPT-User, Claude-User | Whether the relevant service can fetch a page when a user asks it to. |
Blocking GPTBot does not, by itself, block OAI-SearchBot. Blocking every OpenAI user agent may obstruct search fetching you meant to allow. And an open robots.txt file is no victory if your CDN or WAF serves verified crawlers a JavaScript challenge or an HTTP 403. The rule expresses your intent; the response tells you whether a request got through. Decide which access you want, then check whether the requests get through.
Three checks before you buy another audit
I would run these before signing a retainer or paying for another visibility dashboard. They are not a complete citation strategy; they tell you whether the plumbing is working. Run them on a page whose answer matters, not just the homepage. A homepage that loads cleanly cannot tell you what happens to the landing page that holds the useful detail.
- Read the raw page. Run
curl -i -L -A "OAI-SearchBot" https://yoursite.com/your-pageand look for your core product description in the response body. Check the status code and any challenge page first. Missing text might mean client-side rendering, a blocked request, or simply a page that never states the answer. It does not diagnose JavaScript by itself. This request uses a claimed user agent; it does not prove that a verified crawler receives the same response. - Inspect verified requests at the edge. In your CDN or WAF logs, look for legitimate search-crawler requests and their responses. A 403, 429, or challenge tells you something a report of “bot activity” will not: whether the visitor could reach the page at all.
- Check what loads on a click. On a high-value landing page, compare the initial server HTML with what appears after you open an accordion or tab. If the relevant text arrives only after that interaction, put the essential answer somewhere a fetcher can read without performing it.
For a longer implementation path, use our guide to fix technical SEO issues that hurt AI visibility, or work through the 45-minute curl and CDN audit. The useful spend goes toward readable pages, direct answers, and verified crawler access, not a longer queue of decorative fixes. That is also the distinction behind groas: autonomous search work runs continuously within client guardrails, while a named human strategist owns the direction and the commercial result. A dashboard can tell you to check the door. It cannot make an empty page worth reading.
The toggle myth survives because it reduces all of that to allow or block. The harder question is the one I would ask before approving any technical AI-visibility bill: which crawler reached which page, and what could it actually read?
Frequently asked questions
Do bot hits in my server logs mean AI crawlers are reading my site?
No. A user-agent string is a claim made by whoever sent the request, not proof of identity, and scrapers can present themselves as legitimate crawlers. In one 30-day log, 13,405 robots.txt requests came from claimed AI user agents, but only 3,176 came from verified search bots. Verify the visitor, then inspect the actual page request.
Can AI crawlers like GPTBot render JavaScript on my React site?
Do not assume so. An analysis of more than 500 million GPTBot requests found no evidence of client-side JavaScript execution; GPTBot fetched .js files in 11.5% of visits and ClaudeBot in 23.8%, but fetching a script is not the same as running it. If your key text only appears after scripts run, check the initial server response instead of the browser view.
Does adding JSON-LD schema improve my chances of AI citations?
Not on its own. In an Ahrefs study of 1,885 pages that added JSON-LD against 4,000 matched controls, Google AI Overview citations fell by 4.6%, and the changes for ChatGPT (+2.2%) and Google AI Mode (+2.4%) were statistically indistinguishable from zero. The readable page still has to answer the user's question in its body text.
Is an llms.txt file worth adding for AI visibility?
Only marginally. In a 12-week study of 83 domains, OpenAI crawlers requested llms.txt just seven times against 3,990 robots.txt requests, and PerplexityBot requested it zero times; another 90-day benchmark put llms.txt requests at 0.13% of total AI crawler requests. Guessable Markdown page URLs such as /page.md were fetched far more often, and those pages still need to contain the answer.
Are my PPC landing pages ready for AI search if they convert well?
Not necessarily. Pages built for paid clicks often hide specifications in accordions, tabs, or interactive pricing calculators, leaving a non-rendering crawler with a headline, a testimonial, and a form. What matters is whether the text sits in the initial HTML or is inserted only after a click. Keep the facts a crawler needs in the server response.
Can a content-gap report tell me what to write for AI citations?
A content-gap list usually produces familiar SEO phrase-matching work with AI added to the product name, such as repeating enterprise scaling or turnkey integration more often. That is not an answer to a user's actual question. Find the unanswered question and fix the page that buries its explanation, rather than commissioning twenty pages built around suggested phrases.
Can I just block all AI bots in robots.txt with one rule?
No, because AI crawlers serve different functions. Model-training crawlers like GPTBot, search-indexing crawlers like OAI-SearchBot and PerplexityBot, and user-triggered fetchers like ChatGPT-User each require separate decisions; blocking GPTBot does not block OAI-SearchBot. Also check your CDN or WAF, since a 403 or JavaScript challenge can override what robots.txt allows.




