The grievance is the 60-page “AI technical SEO audit” that arrives before anyone checks whether an answer bot can read the site. I was sent one last week: color-coded severity scores, an appendix of schema types no one will ever ship, and a price higher than what I paid for my first car.
On our own site, over 30 days, Claude-User fetched robots.txt 1,194 times during live-answer activity. It fetched the blog index twice. Those counts do not tell me everything it read, and they certainly do not tell me what got cited. They do tell me which technical question kept coming up: is this bot allowed in?
The bot kept checking the door
These were live-answer fetches, not training crawls or background indexing. Claude-User was looking things up while answering questions. The blog index we spend hours polishing received two such fetches that month; the permissions file received 1,194.
I used to tell clients the technical job was polish. I was wrong. The first technical job is access. A beautifully structured page is no help to an answer bot that cannot fetch it. Once the bot gets in, the page still has to say something worth using. robots.txt is not a ranking tactic or a citation button. It is the door.
Google puts a similar boundary around its own AI results: to appear as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible for a snippet, with no additional technical requirements. That does not make good content optional. It makes a 60-page technical wish list harder to defend when the basic eligibility checks have not been done.
The bot names make this less tidy than an audit’s green checkmark suggests. Anthropic runs separate bots: ClaudeBot for training, Claude-SearchBot for search, and Claude-User for live fetches. OpenAI separates GPTBot, OAI-SearchBot, and ChatGPT-User. Blocking a training bot is not the same as blocking the search and user agents. Nor does allowing one automatically allow the others. Each policy needs to name the agent it means.
Yet “allow AI bots” can appear in an audit as a single wildcard line beside a screenshot of a green checkmark. The instruction sounds complete until someone asks which bots the rule covers and whether the edge lets them through. That is not a fix. That is a wish with a badge.
The short version: the bot checked our permissions file 1,194 times. A surprising amount of audit work is advice on repainting a door before anyone tests the lock.
The padding starts on page 14
Page 14 of that PDF told the client to “implement schema on every page.” Forty-two pages later, it was still listing schema types. This is where a reasonable piece of housekeeping gets promoted to the main event.
Ahrefs took 1,885 pages that added JSON-LD and compared them with roughly 4,000 matched controls. The reported changes were AI Mode up 2.4%, ChatGPT up 2.2%, and AI Overviews down 4.6%. The first two were noise, not a demonstrated citation lift. Cited pages were about three times more likely to carry JSON-LD, a number audit vendors can fit nicely on a slide, but the authors read that association as site-quality correlation rather than cause. Every sampled page already had more than 100 AI Overview citations, so the test says nothing about taking a page from zero to its first citation.
None of that says schema is useless. It says the PDF cannot turn “we found more markup to add” into “this will make answer bots cite you.” Google’s guidance, summarized here, treats schema as a floor, not a lever. Reach, parseability, answer position, and fetch speed deserve attention before an appendix of types. I put schema on hundreds of pages by hand back when I believed the checklist. I would rather make sure the crawler can read the answer than sell another round of markup as a citation strategy.
Page 22 prescribed an llms.txt file as if it were a permit office for AI answers. Peec AI looked at roughly 18,000 citations and found six to llms.txt, or 0.03%. Markdown copies drew zero; more than 3,500 citations went to plain HTML instead of a markdown mirror. A separate model across 300,000 domains found that llms.txt added noise and that removing it improved accuracy. I have nothing against the file. I have one on my own site. It can be a convenience for a bot that already reads your HTML. It is not a door, a ramp, or a bribe.
Page 31 was worse: 17 “high severity” findings, each with a red badge and a CVE-style score, none with an owner or a date. Missing alt text on a careers photo. A meta description 12 characters over some tool’s limit. An H1 that appears twice on a template no bot will ever quote. Those may be tasks for someone, someday. A finding without a name next to it and a ship date under it is not engineering. It is inventory.
The padding is the product. A five-item list that says allow the bots, render the page, and move the answer up looks too thin to support a five-figure invoice. Sixty pages with appendices looks like rigor. The client pays for weight and receives weight.

A green checkmark cannot read JavaScript
The second way a site can look open while hiding its answer is JavaScript. A December 2024 crawl study from Vercel and MERJ found that the major AI crawlers it examined did not render JavaScript, including GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, and PerplexityBot. They could fetch JS files without executing them. Gemini on Googlebot infrastructure and Applebot were exceptions.
That distinction is boring until the buying guide exists only after a tab opens. Then it is the entire job. I learned that on an ecommerce account: Google indexed the guide, but the answer text was injected on click. View-source showed what the crawler could see. The bot was not going to click the tab because an audit gave the page a passing grade.
The fix was to put the opening answer in server-rendered HTML, where a fetch could find it. The audit had flagged 40 schema warnings on the template and missed the unreadable words. If the answer lives behind hydration, markup advice comes too late. Check the source before commissioning another taxonomy of warnings.
This is why I object to the order of operations, not to every item in the PDF. Alt text has a job. Schema has a job. A current sitemap has a job. None can substitute for the answer being present in the page the bot receives. “Comprehensive” becomes an expensive word for refusing to choose what matters first.
Five checks before page 55
I would send the vendor five checks before I let them expand the findings list. They are not a promise of citations. They establish whether the pages are accessible, readable, and eligible enough to make the content question worth asking. Run them on the pages you want cited, not just a homepage that happens to pass.
- Read the page with JavaScript off. View-source should contain the answer text, not an empty shell awaiting a click or hydration.
- Find the answer in the first passage. Put a direct answer in the first 120 words of plain HTML; do not make a fetch wade through company history.
- Check canonical and index signals. The canonical, indexability, and snippet eligibility should agree on which page is meant to appear.
- Check the sitemap and page response. Confirm a current lastmod, a 200 response, and an internal path to the URL within three clicks.
- Check each search and user agent in robots.txt and at the CDN or WAF. Verify OAI-SearchBot, Claude-SearchBot, ChatGPT-User, and Claude-User against the edge logs. This is the item vendors skip; a blocked fetch can waste every other fix on that page.
That last check is where the green badge often falls apart. Allowing an agent in robots.txt means little if Cloudflare or another WAF returns a 403. I have seen a green audit checkmark and a 403 in the server log on the same day. Pull a week of logs and filter for the agents you intend to allow. A permissions-file fetch followed by a block is something to investigate at the edge. No fetches at all do not prove a block; they tell you to keep looking rather than declare success from a screenshot.
If you want out of training data but still want search and user agents to reach you, separate those policies. Allow OAI-SearchBot, Claude-SearchBot, and the user agents you want. Disallow GPTBot and ClaudeBot if you do not want their training crawls. Blocking Google-Extended does not hurt normal Google rankings. This is a specific access decision, not a request to “allow AI” and hope the wildcard knows what you meant.
The middle checks lack the glamour of a new file or a severity score. They also expose problems a person can fix. If the answer begins after 800 words about your company history, move it up. If the canonical points away from the page you want found, settle the conflict. If the URL in the sitemap does not return the page you expect, fix the response. Say you run 200 service pages: repairing the shared template is a better first assignment than cataloging 40 schema warnings on pages an answer bot cannot read.
Ask for the fix, not another finding
Here is the question I would put to the vendor: which tool fixes the technical problem that blocks the fetch, and which platform shows the content gaps that keep an accessible page uncited? That is the job groas was built for. It audits what blocks the fetch, fixes what blocks the read, then tracks whether AI answers cite you. Ask to see the work in the page response and the logs before paying for the rest of the PDF. A score cannot stand in for either.

Here is the email I would send back with the audit still open: fix the five checks, show me the log lines, then we can talk about page 14. Show me Claude-User and ChatGPT-User returning 200 on the money pages. Show me view-source with the answer in the first passage. Show me canonicals that agree with the sitemap. Until then, I am not paying for schema inventory and severity confetti.
If the site passes those checks and still gets no citations, stop calling it a plumbing problem. Work on the content and authority: the writing, the mentions, the testing. A bot that can fetch a page is free not to use it. That is an uncomfortable answer for someone whose invoice depends on finding another technical issue, but it is an answer the client can act on.
Keep one habit after the PDF goes away. Once a month, pull the live-fetch lines and count what the bots actually opened. Say you spend $20k a month on content and the log shows 300 permission checks and four article reads. That is a better prompt for investigation than a severity score. Check access. Check the page the bot receives. Check the first paragraph. Do not pay someone to count your H1s while those questions remain unanswered.
The vendor will call the extra pages completeness. I call it billing by the pound. The bot checked the door 1,194 times; the audit sold you a doorstop.

