---
title: "The AI-Visibility Pre-Flight: 27 Checks Before You Pay for ChatGPT Optimization"
description: "Run 27 technical checks for AI visibility, from bot access and raw HTML to sitemaps and verified fetches, before you spend on more content."
url: "https://groas.com/post/the-ai-visibility-technical-pre-flight-2"
image: "https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/0adebcb8-957c-4ce6-bcbe-8f433824caea.png"
published: "2026-10-11T05:39:18.280Z"
modified: "2026-10-11T05:39:18.350Z"
---

October 11, 2026 · 9 min read

# The AI-Visibility Pre-Flight: 27 Checks Before You Pay for ChatGPT Optimization

[Alexander PerelmanHead Of Product @ groas](https://groas.com/author/alexander-perelman)[LinkedIn](https://www.linkedin.com/in/alexander-433793253/)

![A funnel stuffed with printed articles feeds a pipe whose red valve is shut, so an empty glass sits under a dry tap: content can't flow past the block.](https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/0adebcb8-957c-4ce6-bcbe-8f433824caea.png)

In this article

1. [Access: can a live-answer fetcher reach the page?](#access-can-a-live-answer-fetcher-reach-the-page)
2. [Rendering: is the answer in the first HTML response?](#rendering-is-the-answer-in-the-first-html-response)
3. [Discoverability: can a fetcher find the answer URL?](#discoverability-can-a-fetcher-find-the-answer-url)
4. [Page structure: can someone lift the answer in one grab?](#page-structure-can-someone-lift-the-answer-in-one-grab)
5. [Entity and schema: do all versions of the fact agree?](#entity-and-schema-do-all-versions-of-the-fact-agree)
6. [Verification: did a verified fetch reach the fixed page?](#verification-did-a-verified-fetch-reach-the-fixed-page)
7. [Tool or human: who owns the fix?](#tool-or-human-who-owns-the-fix)
8. [Last gate: the edge block you forgot](#last-gate-the-edge-block-you-forgot)

Most vendors will sell you ChatGPT optimization before they check whether a live-answer fetcher can read your site. I would check the plumbing first. In our own edge logs, `robots.txt` was the single most-fetched path on verified Claude-User live-answer requests, ahead of every article and product page we track. Permission is only the first gate: JavaScript can hide the answer, a sitemap can omit it, and an orphaned page can leave a crawler with no route in. Run these 27 checks in order. If you fail one, fix it before you pay for more content.

## Access: can a live-answer fetcher reach the page?

Training crawlers and live-answer fetchers do different jobs. Keep that distinction in mind when you read the rules: [OpenAI separates OAI-SearchBot, ChatGPT-User and GPTBot](https://platform.openai.com/docs/bots), while [Anthropic says its Claude bots honor `robots.txt`](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web). **An Allow rule is not proof of access.** The answer URL also has to survive your edge controls and return a usable page.

![Valves labeled with AI bot names controlling access to a website](https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/951361b3-4cec-493e-9a6b-79ce037304cd.png)

1. **Allow `OAI-SearchBot` on answer paths.** Fetch `/robots.txt` for each relevant host and check whether any applicable `Disallow` covers your money pages. [Blocking OAI-SearchBot keeps a site out of ChatGPT search answers](https://platform.openai.com/docs/bots); blocking a training crawler is a separate decision.
2. Allow `ChatGPT-User` on answer paths. Check its stanza separately rather than assuming the OAI-SearchBot rule covers it. [OpenAI says `robots.txt` rules may not apply to user-initiated ChatGPT-User fetches](https://platform.openai.com/docs/bots), but that is no reason to write a contradictory rule.
3. **Allow `Claude-SearchBot` and `Claude-User` if you want Claude retrieval.** Check the file on every relevant subdomain. Anthropic says these bots honor `robots.txt`, so a rule that looks harmless in a training-crawler audit can close the door on a live answer.
4. **Allow `PerplexityBot` on answer paths if Perplexity matters to you.** Look for a wildcard `Disallow: /` as well as a bot-specific block. The practical split is to [allow search and user-triggered fetchers while handling training crawlers separately](https://github.com/elmohq/elmo/blob/HEAD/packages/docs/content/blog/robots-txt-ai-crawlers.mdx).
5. Put training-crawler decisions in separate stanzas for `GPTBot`, `CCBot` and `Google-Extended`. If you opt out of training, check that `GPTBot` has `Disallow: /` while `OAI-SearchBot` remains allowed. Do not turn one opt-out into a blanket block on every AI bot.
6. **Check for an edge block that overrides `robots.txt`.** Inspect Cloudflare bot settings, WAF rules and edge responses for the target URL; a normal browser request alone will not prove a bot can get through. [Cloudflare’s Block AI bots setting can stop requests at the edge](https://blog.cloudflare.com/declaring-your-aindependence-block-ai-bots-scrapers-and-crawlers-with-a-single-click/) even when `robots.txt` says Allow.
7. Make each key answer URL return 200 without a redirect chain. Request the exact URL you want cited, then test its trailing-slash variant. A clean rule for the homepage does not rescue a product page that lands on a 404 after two hops.
8. Rule out geo- and IP-based blocks on verified fetchers. Check edge and server logs for 403s on your answer URLs; do not infer bot access from your own location. Before trusting a `ChatGPT-User` log entry, [check its IP against OpenAI’s published JSON](https://openai.com/chatgpt-user.json). For rules that preserve the crawler split, use [The AI Crawler Swipe File](https://groas.com/post/the-ai-crawler-swipe-file-copy-paste-rob).

## Rendering: is the answer in the first HTML response?

**View source, not Inspect Element.** [Major AI crawlers fetch JavaScript files without executing them](https://vercel.com/blog/the-rise-of-the-ai-crawler). The draft’s numbers are a useful warning: ChatGPT fetched JS on only 11.5% of requests and Claude on 23.84%. What a deck calls a “dynamic experience” may be an empty div to the fetcher. Use `curl` or view-source for these checks; what appears after the browser runs scripts is not the test.

9. **Put the core answer, price and specs in server-rendered HTML.** Search the raw response for the complete answer paragraph. If the source contains only a placeholder and the answer appears after JavaScript runs, the page fails, however polished it looks in a browser.
10. **Include FAQ, tab and accordion answers in the HTML.** Check the source for every answer pane, not just the default open tab. A crawler will not click through the interface to find the price or condition you tucked behind it.
11. Keep answer images and tables out of a JavaScript-only lazy-load gate. Disable JS and request the page again. If the relevant content disappears, render that answer block on the server; this check does not require rebuilding the whole site.
12. Keep key pages under roughly 150KB of raw HTML before images. Check response size and whether the answer sits near the start rather than after a mountain of builder markup. If this fails, run [the 45-minute site-readability check](https://groas.com/post/can-chatgpt-even-read-your-site-a-45-min) before writing another answer page.

## Discoverability: can a fetcher find the answer URL?

**Allowed does not mean discovered.** A fetcher can have permission to visit every page and still miss an answer that appears in no sitemap or internal link. Check the route to the page, not just its response once you already know the URL.

13. List answer pages in a working XML sitemap. Request the sitemap, confirm it returns 200, and find each target URL in it. If the only copy of a useful fact sits in an orphaned PDF, give the crawler a discoverable page for that answer.
14. **Make `lastmod` honest or omit it.** Sample five URLs and compare their dates with the last significant edits. [Google uses `lastmod` when it is consistently accurate and ignores `priority` and `changefreq`](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap). Stamping every URL with today’s date is not a freshness strategy.
15. **Link to each answer page within three clicks of the homepage or main hub.** Use a crawl-depth report or walk the path yourself. A sitemap entry helps, but an orphaned page still gives visitors and crawlers no obvious route from the rest of the site.
16. Point each canonical tag at the answer URL you want cited. Inspect `rel=canonical` on the target page and check that it self-references. Do not make the visible answer page nominate a different URL while you try to earn citations for this one.
17. Treat `llms.txt` as optional garnish, not a repair. Its presence does not fix blocked requests, orphaned pages or missing HTML answers. [Google says it does not use `llms.txt`-style files for Search visibility and rankings](https://developers.google.com/search/docs/appearance/structured-data/faqpage). Fix the route to the page first.

## Page structure: can someone lift the answer in one grab?

**Put the extractable fact where a retriever finds it.** I used to bury the lede for readability. I was wrong. If a reader needs four paragraphs and an image of a spec table to work out the answer, a fetcher has the same problem.

18. **Answer the main question in the first 80 to 100 words.** Read that passage alone: does it give the price, definition or first step without a windup? If not, move the direct answer up and let the explanation follow it.
19. Give each main question its own `H2`, with the answer immediately below. Read the headings as queries, then read only the first sentence under each. If those sentences dodge the questions, rewrite them before adding more headings.
20. **Put tables and specs in selectable text, not only images or embedded PDFs.** Try to highlight a spec row with your cursor and find it in the HTML response. If the fact exists only as pixels, do not expect a fetcher to quote it reliably.
21. Show a visible date and author on advice pages. Check that an updated date in the content fits the shelf life of the claim it accompanies. A current-looking page with stale prices or steps is still a stale answer.

## Entity and schema: do all versions of the fact agree?

**Markup should repeat what a human can see.** Schema does not earn a citation by itself. If the visible price, `Product` markup and FAQ give three prices, you have made the answer harder to trust, not easier to retrieve.

22. **Keep one consistent `Organization` identity across pages.** Compare name, URL and contact details in the markup and visible site copy, especially after a rebrand. A validator can catch mismatched fields; it cannot decide which old name your business meant to keep.
23. Match any `Product` or `Service` and `Offer` price to the visible price. Compare rendered copy with markup on five money pages. If the two disagree, settle the real price before changing the schema.
24. **Make FAQ markup mirror visible questions and answers.** Check that every marked-up Q\&A appears on the page, with the same answer a reader gets. Hidden markup-only questions do not repair a thin page.
25. Keep entity facts consistent across the site: hours, locations, plan names and specs. Search for old prices and product names, including in PDFs. One stale copy can muddy retrieval even when the current money page is correct.

![Printed answer page with its first paragraph and spec table highlighted](https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/22a9f409-0cdc-484b-88d6-6c259d72ce2a.png)

## Verification: did a verified fetch reach the fixed page?

**Logs first, prompts second.** A tracked prompt score can move without a fetcher touching the page you changed. Check the request evidence, then check the answer. I would not buy another rewrite because a dashboard line moved.

26. **Find verified fetcher hits on answer URLs, not only `robots.txt`.** Filter the last 14 days for `OAI-SearchBot`, `ChatGPT-User`, `Claude-SearchBot`, `Claude-User` and `PerplexityBot`, then inspect requests to the target pages. A User-Agent string alone proves nothing; [verify `ChatGPT-User` IPs against OpenAI’s published JSON](https://openai.com/chatgpt-user.json).
27. Spot-check three live prompts after each fix batch and record the cited URL. Keep the question and account the same, note the before-and-after dates, and check whether the fixed URL is cited with the right price or steps. If logs show no fetch, wait and check access again before rewriting.

## Tool or human: who owns the fix?

**Automate the repeated checks; keep permission and truth with a person.** That is the useful split if you want to avoid turning this into a six-week agency project. groas’s model puts continuous execution inside guardrails and gives a named strategist ownership of the decisions a script should not make.

- **Tool: checks 1–5 and 13–16.** Check crawler-rule splits, sitemap entries, `lastmod`, canonicals and link depth repeatedly. Flag a failure before another batch of content inherits it.
- **Tool: checks 9–12 and 18–25.** Compare raw HTML with the intended answer, flag buried or image-trapped facts, and check structured data against visible copy after changes go live.
- **Human: checks 6–8.** Approve changes to edge security, WAF rules, redirects and geo-blocks. A tool can expose a 403; the business has to decide which access it permits.
- **Human: checks 26–27 and disputed facts.** Confirm the real price, hours or specs when pages disagree. Then sign off on verified fetches and citation checks before buying more words.

## Last gate: the edge block you forgot

- **Check 6 is the one I would revisit before approving a content budget.** Say you spend $20k a year on AEO content while Cloudflare returns 403 to a live-answer fetcher. Your `robots.txt` can say Allow, your sitemap can list the page, and your copy can be excellent. The request still stops at the edge. That is what the skipped check costs: you pay for answers the fetcher never reads. Fix the valve before you buy more water.

## Frequently Asked Questions

### Why should I check robots.txt and bot access before paying for ChatGPT optimization content?

Because permission is the first gate: robots.txt was the single most-fetched path on verified Claude-User live-answer requests in edge logs, and if a fetcher cannot reach your pages, content you pay for is never read. Check bot rules, edge controls and raw page delivery before buying more content.

### My robots.txt allows AI bots, so am I guaranteed they can fetch my pages?

No. An Allow rule is not proof of access. Cloudflare's Block AI bots setting can stop requests at the edge even when robots.txt says Allow, so you also need to check Cloudflare bot settings, WAF rules and edge responses, and confirm each key answer URL returns 200 without a redirect chain or geo/IP block.

### Does ChatGPT run JavaScript when it reads my website pages?

Usually not. Major AI crawlers fetch JavaScript files without executing them, and ChatGPT fetched JS on only 11.5% of requests and Claude on 23.84% in the draft's numbers. Put the core answer, price and specs in server-rendered HTML, and use curl or view-source rather than Inspect Element to test.

### Can AI crawlers read answers hidden behind tabs, accordions or lazy-loaded content?

Not reliably if those answers only appear after JavaScript runs. Check the raw HTML source for every answer pane, and disable JS to see whether content disappears; if it does, render that answer block on the server. You do not need to rebuild the whole site to fix a single block.

### If a fetcher is allowed to visit my pages, will it automatically find my answer pages?

No, allowed does not mean discovered. A fetcher can have permission for every page and still miss an answer that appears in no XML sitemap or internal link. List answer pages in a working sitemap, keep lastmod honest, and link each answer page within three clicks of the homepage so it is not orphaned.

### Where on the page should the direct answer to a question appear so AI fetchers can quote it?

In the first 80 to 100 words, giving the price, definition or first step without a windup. Also give each main question its own H2 with the answer immediately below, and put tables and specs in selectable text, since a fact that exists only as pixels in an image or embedded PDF cannot be quoted reliably.

### Does adding schema markup automatically earn citations from AI assistants?

No, schema does not earn a citation by itself. Markup should repeat what a human can see: the Organization identity, Product or Offer prices and FAQ Q\&As must match the visible copy. If the visible price, markup and FAQ give three different prices, the answer becomes harder to trust, not easier to retrieve.

### How do I verify that AI fetchers actually reached the page I fixed?

Check logs first, prompts second. Filter the last 14 days for OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User and PerplexityBot requests to the target pages, verifying ChatGPT-User IPs against OpenAI's published JSON. Then spot-check three live prompts with the same question and account, recording the cited URL and before-and-after dates.

## Pay For Results, Not For Hours

Businesses buy the outcome, agencies resell it, and groas answers for it either way.

[See If You Qualify](https://groas.typeform.com/to/xC1bQNUT)

## Structured data

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://graph.groas.com/entity/groas","name":"groas","url":"https://groas.com/","logo":{"@type":"ImageObject","url":"https://groas.com/icon-512.png","width":512,"height":512},"sameAs":["https://www.linkedin.com/company/groas/"]},{"@type":"WebSite","@id":"https://groas.com/#website","name":"groas","url":"https://groas.com/","inLanguage":"en-US","publisher":{"@id":"https://graph.groas.com/entity/groas"}},{"@type":"WebPage","@id":"https://groas.com/post/the-ai-visibility-technical-pre-flight-2#webpage","url":"https://groas.com/post/the-ai-visibility-technical-pre-flight-2","name":"The AI-Visibility Pre-Flight: 27 Checks Before You Pay for ChatGPT Optimization","description":"Run 27 technical checks for AI visibility, from bot access and raw HTML to sitemaps and verified fetches, before you spend on more content.","inLanguage":"en-US","isPartOf":{"@id":"https://groas.com/#website"},"about":{"@id":"https://graph.groas.com/entity/groas"},"primaryImageOfPage":{"@type":"ImageObject","url":"https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/0adebcb8-957c-4ce6-bcbe-8f433824caea.png"},"breadcrumb":{"@id":"https://groas.com/post/the-ai-visibility-technical-pre-flight-2#breadcrumb"},"dateModified":"2026-10-11T05:39:18.350Z","mainEntity":{"@id":"https://groas.com/post/the-ai-visibility-technical-pre-flight-2#article"}},{"@type":"BreadcrumbList","@id":"https://groas.com/post/the-ai-visibility-technical-pre-flight-2#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://groas.com/"},{"@type":"ListItem","position":2,"name":"Blog","item":"https://groas.com/blog"},{"@type":"ListItem","position":3,"name":"The AI-Visibility Pre-Flight: 27 Checks Before You Pay for ChatGPT Optimization","item":"https://groas.com/post/the-ai-visibility-technical-pre-flight-2"}]},{"@type":"BlogPosting","@id":"https://groas.com/post/the-ai-visibility-technical-pre-flight-2#article","url":"https://groas.com/post/the-ai-visibility-technical-pre-flight-2","mainEntityOfPage":{"@id":"https://groas.com/post/the-ai-visibility-technical-pre-flight-2#webpage"},"isPartOf":{"@id":"https://groas.com/blog#blog"},"headline":"The AI-Visibility Pre-Flight: 27 Checks Before You Pay for ChatGPT Optimization","description":"Run 27 technical checks for AI visibility, from bot access and raw HTML to sitemaps and verified fetches, before you spend on more content.","datePublished":"2026-10-11T05:39:18.280Z","dateModified":"2026-10-11T05:39:18.350Z","image":{"@type":"ImageObject","url":"https://pub-87da24ecbbfc4c3bad6875f3aa013712.r2.dev/generated-images/0adebcb8-957c-4ce6-bcbe-8f433824caea.png"},"author":{"@id":"https://groas.com/author/alexander-perelman#person"},"publisher":{"@id":"https://graph.groas.com/entity/groas"},"keywords":"answer, check, page, before, checks, url, checks before, chatgpt optimization","wordCount":1825,"inLanguage":"en-US"},{"@type":"Person","@id":"https://groas.com/author/alexander-perelman#person","name":"Alexander Perelman","jobTitle":"Head Of Product @ groas","description":"Ex Goldman Sachs and Ex Stanford Computer Science","image":"https://groas.com/media/blog/4f17cac3a81acc224cc1b2cfab8f94bbc295a3089d13e8421125771ef8d4064f.jpg","worksFor":{"@id":"https://graph.groas.com/entity/groas"},"sameAs":["https://www.linkedin.com/in/alexander-433793253/"]},{"@type":"FAQPage","@id":"https://groas.com/post/the-ai-visibility-technical-pre-flight-2#faq","mainEntity":[{"@type":"Question","name":"Why should I check robots.txt and bot access before paying for ChatGPT optimization content?","acceptedAnswer":{"@type":"Answer","text":"Because permission is the first gate: robots.txt was the single most-fetched path on verified Claude-User live-answer requests in edge logs, and if a fetcher cannot reach your pages, content you pay for is never read. Check bot rules, edge controls and raw page delivery before buying more content."}},{"@type":"Question","name":"My robots.txt allows AI bots, so am I guaranteed they can fetch my pages?","acceptedAnswer":{"@type":"Answer","text":"No. An Allow rule is not proof of access. Cloudflare's Block AI bots setting can stop requests at the edge even when robots.txt says Allow, so you also need to check Cloudflare bot settings, WAF rules and edge responses, and confirm each key answer URL returns 200 without a redirect chain or geo/IP block."}},{"@type":"Question","name":"Does ChatGPT run JavaScript when it reads my website pages?","acceptedAnswer":{"@type":"Answer","text":"Usually not. Major AI crawlers fetch JavaScript files without executing them, and ChatGPT fetched JS on only 11.5% of requests and Claude on 23.84% in the draft's numbers. Put the core answer, price and specs in server-rendered HTML, and use curl or view-source rather than Inspect Element to test."}},{"@type":"Question","name":"Can AI crawlers read answers hidden behind tabs, accordions or lazy-loaded content?","acceptedAnswer":{"@type":"Answer","text":"Not reliably if those answers only appear after JavaScript runs. Check the raw HTML source for every answer pane, and disable JS to see whether content disappears; if it does, render that answer block on the server. You do not need to rebuild the whole site to fix a single block."}},{"@type":"Question","name":"If a fetcher is allowed to visit my pages, will it automatically find my answer pages?","acceptedAnswer":{"@type":"Answer","text":"No, allowed does not mean discovered. A fetcher can have permission for every page and still miss an answer that appears in no XML sitemap or internal link. List answer pages in a working sitemap, keep lastmod honest, and link each answer page within three clicks of the homepage so it is not orphaned."}},{"@type":"Question","name":"Where on the page should the direct answer to a question appear so AI fetchers can quote it?","acceptedAnswer":{"@type":"Answer","text":"In the first 80 to 100 words, giving the price, definition or first step without a windup. Also give each main question its own H2 with the answer immediately below, and put tables and specs in selectable text, since a fact that exists only as pixels in an image or embedded PDF cannot be quoted reliably."}},{"@type":"Question","name":"Does adding schema markup automatically earn citations from AI assistants?","acceptedAnswer":{"@type":"Answer","text":"No, schema does not earn a citation by itself. Markup should repeat what a human can see: the Organization identity, Product or Offer prices and FAQ Q&As must match the visible copy. If the visible price, markup and FAQ give three different prices, the answer becomes harder to trust, not easier to retrieve."}},{"@type":"Question","name":"How do I verify that AI fetchers actually reached the page I fixed?","acceptedAnswer":{"@type":"Answer","text":"Check logs first, prompts second. Filter the last 14 days for OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User and PerplexityBot requests to the target pages, verifying ChatGPT-User IPs against OpenAI's published JSON. Then spot-check three live prompts with the same question and account, recording the cited URL and before-and-after dates."}}]}]}
```
