Most vendors will sell you ChatGPT optimization before they check whether a live-answer fetcher can read your site. I would check the plumbing first. In our own edge logs, robots.txt was the single most-fetched path on verified Claude-User live-answer requests, ahead of every article and product page we track. Permission is only the first gate: JavaScript can hide the answer, a sitemap can omit it, and an orphaned page can leave a crawler with no route in. Run these 27 checks in order. If you fail one, fix it before you pay for more content.
Access: can a live-answer fetcher reach the page?
Training crawlers and live-answer fetchers do different jobs. Keep that distinction in mind when you read the rules: OpenAI separates OAI-SearchBot, ChatGPT-User and GPTBot, while Anthropic says its Claude bots honor robots.txt. An Allow rule is not proof of access. The answer URL also has to survive your edge controls and return a usable page.

- Allow
OAI-SearchBoton answer paths. Fetch/robots.txtfor each relevant host and check whether any applicableDisallowcovers your money pages. Blocking OAI-SearchBot keeps a site out of ChatGPT search answers; blocking a training crawler is a separate decision. - Allow
ChatGPT-Useron answer paths. Check its stanza separately rather than assuming the OAI-SearchBot rule covers it. OpenAI saysrobots.txtrules may not apply to user-initiated ChatGPT-User fetches, but that is no reason to write a contradictory rule. - Allow
Claude-SearchBotandClaude-Userif you want Claude retrieval. Check the file on every relevant subdomain. Anthropic says these bots honorrobots.txt, so a rule that looks harmless in a training-crawler audit can close the door on a live answer. - Allow
PerplexityBoton answer paths if Perplexity matters to you. Look for a wildcardDisallow: /as well as a bot-specific block. The practical split is to allow search and user-triggered fetchers while handling training crawlers separately. - Put training-crawler decisions in separate stanzas for
GPTBot,CCBotandGoogle-Extended. If you opt out of training, check thatGPTBothasDisallow: /whileOAI-SearchBotremains allowed. Do not turn one opt-out into a blanket block on every AI bot. - Check for an edge block that overrides
robots.txt. Inspect Cloudflare bot settings, WAF rules and edge responses for the target URL; a normal browser request alone will not prove a bot can get through. Cloudflare’s Block AI bots setting can stop requests at the edge even whenrobots.txtsays Allow. - Make each key answer URL return 200 without a redirect chain. Request the exact URL you want cited, then test its trailing-slash variant. A clean rule for the homepage does not rescue a product page that lands on a 404 after two hops.
- Rule out geo- and IP-based blocks on verified fetchers. Check edge and server logs for 403s on your answer URLs; do not infer bot access from your own location. Before trusting a
ChatGPT-Userlog entry, check its IP against OpenAI’s published JSON. For rules that preserve the crawler split, use The AI Crawler Swipe File.
Rendering: is the answer in the first HTML response?
View source, not Inspect Element. Major AI crawlers fetch JavaScript files without executing them. The draft’s numbers are a useful warning: ChatGPT fetched JS on only 11.5% of requests and Claude on 23.84%. What a deck calls a “dynamic experience” may be an empty div to the fetcher. Use curl or view-source for these checks; what appears after the browser runs scripts is not the test.
- Put the core answer, price and specs in server-rendered HTML. Search the raw response for the complete answer paragraph. If the source contains only a placeholder and the answer appears after JavaScript runs, the page fails, however polished it looks in a browser.
- Include FAQ, tab and accordion answers in the HTML. Check the source for every answer pane, not just the default open tab. A crawler will not click through the interface to find the price or condition you tucked behind it.
- Keep answer images and tables out of a JavaScript-only lazy-load gate. Disable JS and request the page again. If the relevant content disappears, render that answer block on the server; this check does not require rebuilding the whole site.
- Keep key pages under roughly 150KB of raw HTML before images. Check response size and whether the answer sits near the start rather than after a mountain of builder markup. If this fails, run the 45-minute site-readability check before writing another answer page.
Discoverability: can a fetcher find the answer URL?
Allowed does not mean discovered. A fetcher can have permission to visit every page and still miss an answer that appears in no sitemap or internal link. Check the route to the page, not just its response once you already know the URL.
- List answer pages in a working XML sitemap. Request the sitemap, confirm it returns 200, and find each target URL in it. If the only copy of a useful fact sits in an orphaned PDF, give the crawler a discoverable page for that answer.
- Make
lastmodhonest or omit it. Sample five URLs and compare their dates with the last significant edits. Google useslastmodwhen it is consistently accurate and ignorespriorityandchangefreq. Stamping every URL with today’s date is not a freshness strategy. - Link to each answer page within three clicks of the homepage or main hub. Use a crawl-depth report or walk the path yourself. A sitemap entry helps, but an orphaned page still gives visitors and crawlers no obvious route from the rest of the site.
- Point each canonical tag at the answer URL you want cited. Inspect
rel=canonicalon the target page and check that it self-references. Do not make the visible answer page nominate a different URL while you try to earn citations for this one. - Treat
llms.txtas optional garnish, not a repair. Its presence does not fix blocked requests, orphaned pages or missing HTML answers. Google says it does not usellms.txt-style files for Search visibility and rankings. Fix the route to the page first.
Page structure: can someone lift the answer in one grab?
Put the extractable fact where a retriever finds it. I used to bury the lede for readability. I was wrong. If a reader needs four paragraphs and an image of a spec table to work out the answer, a fetcher has the same problem.
- Answer the main question in the first 80 to 100 words. Read that passage alone: does it give the price, definition or first step without a windup? If not, move the direct answer up and let the explanation follow it.
- Give each main question its own
H2, with the answer immediately below. Read the headings as queries, then read only the first sentence under each. If those sentences dodge the questions, rewrite them before adding more headings. - Put tables and specs in selectable text, not only images or embedded PDFs. Try to highlight a spec row with your cursor and find it in the HTML response. If the fact exists only as pixels, do not expect a fetcher to quote it reliably.
- Show a visible date and author on advice pages. Check that an updated date in the content fits the shelf life of the claim it accompanies. A current-looking page with stale prices or steps is still a stale answer.
Entity and schema: do all versions of the fact agree?
Markup should repeat what a human can see. Schema does not earn a citation by itself. If the visible price, Product markup and FAQ give three prices, you have made the answer harder to trust, not easier to retrieve.
- Keep one consistent
Organizationidentity across pages. Compare name, URL and contact details in the markup and visible site copy, especially after a rebrand. A validator can catch mismatched fields; it cannot decide which old name your business meant to keep. - Match any
ProductorServiceandOfferprice to the visible price. Compare rendered copy with markup on five money pages. If the two disagree, settle the real price before changing the schema. - Make FAQ markup mirror visible questions and answers. Check that every marked-up Q&A appears on the page, with the same answer a reader gets. Hidden markup-only questions do not repair a thin page.
- Keep entity facts consistent across the site: hours, locations, plan names and specs. Search for old prices and product names, including in PDFs. One stale copy can muddy retrieval even when the current money page is correct.

Verification: did a verified fetch reach the fixed page?
Logs first, prompts second. A tracked prompt score can move without a fetcher touching the page you changed. Check the request evidence, then check the answer. I would not buy another rewrite because a dashboard line moved.
- Find verified fetcher hits on answer URLs, not only
robots.txt. Filter the last 14 days forOAI-SearchBot,ChatGPT-User,Claude-SearchBot,Claude-UserandPerplexityBot, then inspect requests to the target pages. A User-Agent string alone proves nothing; verifyChatGPT-UserIPs against OpenAI’s published JSON. - Spot-check three live prompts after each fix batch and record the cited URL. Keep the question and account the same, note the before-and-after dates, and check whether the fixed URL is cited with the right price or steps. If logs show no fetch, wait and check access again before rewriting.
Tool or human: who owns the fix?
Automate the repeated checks; keep permission and truth with a person. That is the useful split if you want to avoid turning this into a six-week agency project. groas’s model puts continuous execution inside guardrails and gives a named strategist ownership of the decisions a script should not make.
- Tool: checks 1–5 and 13–16. Check crawler-rule splits, sitemap entries,
lastmod, canonicals and link depth repeatedly. Flag a failure before another batch of content inherits it. - Tool: checks 9–12 and 18–25. Compare raw HTML with the intended answer, flag buried or image-trapped facts, and check structured data against visible copy after changes go live.
- Human: checks 6–8. Approve changes to edge security, WAF rules, redirects and geo-blocks. A tool can expose a 403; the business has to decide which access it permits.
- Human: checks 26–27 and disputed facts. Confirm the real price, hours or specs when pages disagree. Then sign off on verified fetches and citation checks before buying more words.
Last gate: the edge block you forgot
- Check 6 is the one I would revisit before approving a content budget. Say you spend $20k a year on AEO content while Cloudflare returns 403 to a live-answer fetcher. Your
robots.txtcan say Allow, your sitemap can list the page, and your copy can be excellent. The request still stops at the edge. That is what the skipped check costs: you pay for answers the fetcher never reads. Fix the valve before you buy more water.

