---
title: "The AI Visibility Blocks Your Site Forgot: A Timeline from GPTBot to CDN Defaults"
description: "Your site may be missing from AI answers because someone made a perfectly sensible decision in 2023 and nobody checked it again."
url: "https://groas.com/post/from-gptbot-to-default-blocking-a-dated"
image: "https://groas.com/media/blog/8e8379bcc6e1091943afb11a15fa233040f034f9b4ceb47cd11526669070d3cd.png"
published: "2026-10-04T05:43:25.149Z"
modified: "2026-10-04T05:43:26.253Z"
---

[Google Ads News](https://groas.com/category/google-ads-news) · October 4, 2026 · 9 min read

# The AI Visibility Blocks Your Site Forgot: A Timeline from GPTBot to CDN Defaults

[DavidFounder & CEO @ groas](https://groas.com/author/david)

![Cover image for: The AI Visibility Blocks Your Site Forgot: A Timeline from GPTBot to CDN Defaults](https://groas.com/media/blog/8e8379bcc6e1091943afb11a15fa233040f034f9b4ceb47cd11526669070d3cd.png)

In this article

1. [Aug 2023: GPTBot gives owners a training opt-out](#aug-2023-gptbot-gives-owners-a-training-opt-out)
2. [Sep 2023: Google-Extended adds another easily misread rule](#sep-2023-google-extended-adds-another-easily-misread-rule)
3. [2024: Answer bots get names of their own](#2024-answer-bots-get-names-of-their-own)
4. [Jul 2024: Cloudflare puts an AI block below robots.txt](#jul-2024-cloudflare-puts-an-ai-block-below-robotstxt)
5. [Sep 2024: llms.txt offers a shortcut that cannot open a blocked page](#sep-2024-llmstxt-offers-a-shortcut-that-cannot-open-a-blocked-page)
6. [Jul 2025: Blocking becomes the starting setting for new domains](#jul-2025-blocking-becomes-the-starting-setting-for-new-domains)
7. [2025: User-agent rules collide with fetches and firewalls](#2025-user-agent-rules-collide-with-fetches-and-firewalls)
8. [Oct 2026: Date the fossil before you touch the content](#oct-2026-date-the-fossil-before-you-touch-the-content)

Your site may be missing from AI answers because someone made a perfectly sensible decision in 2023 and nobody checked it again. The problem is not always a new technical SEO bug. Often, it is a fossil: a robots.txt rule, CDN toggle or bot block left over from the 2023–2025 scramble to keep content out of AI training.

The year matters. A GPTBot block from 2023 expresses a training preference; a copied block list from 2024 may shut out answer bots; a CDN setting from 2025 can make your robots.txt file irrelevant. Date the setting before you change it. Otherwise, you risk deleting the rule you meant to keep while leaving the actual block in place.

## Aug 2023: GPTBot gives owners a training opt-out

On Aug 7, 2023, OpenAI [documented GPTBot as a web crawler for training data with its own robots.txt token](https://platform.openai.com/docs/bots). The choice was straightforward: add `User-agent: GPTBot` and `Disallow: /` if you did not want GPTBot crawling your content for training. It was a low-friction way to make that choice, so the rule spread.

That decision still turns up in audits, often treated as the explanation for every missing ChatGPT citation. It is not. **A `Disallow` for GPTBot alone does not block ChatGPT search answers.** OpenAI later documented separate bots for separate jobs. GPTBot remained the training crawler; answer-related access got other names.

If your file only blocks GPTBot, keep the rule if you still want the training opt-out. Then keep looking. This is the oldest fossil in the stack, but it is not necessarily the one costing you citations.

## Sep 2023: Google-Extended adds another easily misread rule

Late September 2023 brought [Google-Extended, a robots.txt control token for Bard and Vertex AI training rather than a separate crawler or a Google Search indexing control](https://blog.getadmiral.com/how-to-opt-out-of-ai-training-bots-by-google-bard-and-openai-chatgpt). Owners put it beside GPTBot in the same file for much the same reason: they wanted a say in how their content was used for AI.

The file location encouraged a bad inference. Because both controls sat in robots.txt, it was easy to assume both governed visibility in search results and AI answers. **A Google-Extended block is not a Google Search exclusion.** Do not remove a training preference simply because a page is absent from AI Overviews, and do not expect removing that preference to repair a crawler or rendering problem elsewhere.

The 2023 rules show why dating matters: two entries that look like visibility controls may instead record decisions about training. Keep the decisions you intended. Troubleshoot answer access separately.

## 2024: Answer bots get names of their own

OpenAI’s [documented crawler roles](https://platform.openai.com/docs/bots) now separate GPTBot for training, OAI-SearchBot for the ChatGPT search index, and ChatGPT-User for on-demand fetches. The distinction changes the audit. A GPTBot block can coexist with access for OAI-SearchBot; blocking OAI-SearchBot, by contrast, keeps a site out of the ChatGPT search path it serves.

I used to treat a GPTBot block as the whole AI-access decision. I was wrong. **The bot’s job, not the word “AI” on a block list, tells you what the rule does.** A copied list that disallows every newly named bot may protect a training preference and cut off answers in the same edit.

The distinction extends beyond OpenAI. Anthropic [documents ClaudeBot for training, Claude-SearchBot for search, and Claude-User for user-directed fetches](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler). Perplexity [describes PerplexityBot as serving its search results and Perplexity-User as fetching in response to user requests](https://docs.perplexity.ai/docs/resources/perplexity-crawlers). Those names are not interchangeable, and neither are the consequences of blocking them.

This is where an old intention became a new problem. A 2023 file mentioning only GPTBot probably does not explain a missing answer. A later rule that also catches OAI-SearchBot, Claude-SearchBot or PerplexityBot might. Read the actual user-agent groups before you rewrite a page.

## Jul 2024: Cloudflare puts an AI block below robots.txt

In July 2024, Cloudflare [released its one-click Block AI Scrapers and Crawlers control under Security > Bots, including on free plans](https://blog.cloudflare.com/declaring-your-aindependence-block-ai-bots-scrapers-and-crawlers-with-a-single-click/). It updated as Cloudflare identified more bots. For an owner trying to limit scraping, a toggle was considerably easier than maintaining a list of names in robots.txt.

It also moved the decision to a different layer. **A robots.txt `Allow` cannot undo a network-layer block.** An owner can inspect a tidy file, see permission for an answer bot and still have the CDN turn that bot away before it reaches the page. That is why a robots-only audit can look reassuring while live fetches fail.

Cloudflare reported at launch that GPTBot had reached 35.46% of its sites and ClaudeBot 11.17%, while 2.98% of top sites were blocking. The toggle addressed a real concern about unwanted crawling. But if the goal later changes from blocking crawlers to appearing in AI answers, revisit the toggle rather than assuming your original robots.txt choices still describe access. Check the CDN alongside the file.

## Sep 2024: llms.txt offers a shortcut that cannot open a blocked page

On Sept 3, 2024, Jeremy Howard and Answer.AI [proposed _llms.txt_: a Markdown file at `/llms.txt` with a summary and links meant to make useful site content easier for models to find](https://llmstxt.org/). It was a community proposal, not an IETF standard. The attraction was obvious: give a model a clean route through the site instead of asking it to sort through navigation and JavaScript.

Then came the familiar shortcut. A site could publish an auto-generated `/llms.txt`, declare itself ready for AI search and leave its robots.txt rules and CDN block untouched. Some files pointed to a few old posts while the pages that mattered most remained inaccessible. That is not an access strategy. It is a table of contents outside a locked building.

**An `/llms.txt` file does not override robots.txt or a firewall.** If you have one, make sure its links reflect the pages you want found, then test whether the relevant bots can fetch those pages. If you do not have one, its absence is not, by itself, an explanation for invisibility. Audit the door before polishing the sign.

## Jul 2025: Blocking becomes the starting setting for new domains

On July 1, 2025, Cloudflare [made blocking AI crawlers the default for new domains unless owners opted in or used Pay Per Crawl](https://blog.cloudflare.com/content-independence-day-no-ai-crawl-without-compensation/). A setting that previously required a click could now be in place when a domain was added. The practical consequence is easy to miss during a launch or migration: the team may never have made an explicit decision to block answer-related access.

That makes a sudden loss of AI-answer visibility a different diagnostic problem. Before changing copy, inspect the CDN setting on the domain that is actually serving the pages. Before editing a bot rule, confirm a fetch reaches the server. **A clean robots.txt file is not proof that a crawler can get through Cloudflare.**

This entry is also a useful reminder that not every fossil was added by a person in a hurry. Some arrived as defaults. If the timing of a visibility change lines up with a new domain or a migration, check the settings that came with it.

## 2025: User-agent rules collide with fetches and firewalls

The bot names became easier to confuse just as access decisions moved deeper into the stack. OpenAI says [ChatGPT-User fetches are user-initiated, robots.txt rules may not apply to them, and this bot is not the control for ChatGPT Search inclusion](https://platform.openai.com/docs/bots). Perplexity likewise says [Perplexity-User generally ignores robots.txt when a user requests a page and provides IP information for WAF configuration](https://docs.perplexity.ai/docs/resources/perplexity-crawlers).

So a robots.txt `Allow` is not a firewall pass. Nor does a user-agent string, on its own, establish who sent a request. **Check the access decision at the layer making it.** If a WAF blocks a fetch, changing the wording of a robots.txt group will not make the page load. If robots.txt disallows a search bot, an IP allow-list does not express the permission you meant to give in that file.

There was a reason security teams grew cautious. Following Wired’s June 2024 reporting, [Cloudflare’s 2025 tests described Perplexity fetches from disallowed sites using undisclosed IPs and rotated user-agents, and Cloudflare de-listed it as a verified bot](https://www.techrepublic.com/article/news-cloudflare-perplexity-crawling-violations/). That dispute does not make a blanket block on every AI-named bot a precise policy. It makes verification more important.

Separate the decisions: which crawlers may index or fetch your pages, and which requests your firewall will admit. Compare relevant requests with published bot IP information where available, inspect the WAF result, and test the page itself. A user-agent list alone cannot tell you what happened.

## Oct 2026: Date the fossil before you touch the content

The pattern is clearer from here. Owners used named training opt-outs in 2023. In 2024, answer-related bots had distinct names, while copied block lists and CDN controls made it easier to deny them access. By 2025, a new domain could begin with an AI crawler block at the CDN. None of those decisions requires a ranking change to make a page disappear from an answer path. A fetch that never completes is enough.

Here is the audit I run before changing content. **Work from policy to live response:** identify what each rule was meant to control, check every layer that can deny the request, then confirm what the bot receives. The dates are clues, not permission to delete every old setting.

1. **Check `User-agent: GPTBot` and `Disallow: /` (Aug 2023).** Keep the training opt-out if you meant it; do not mistake it for an answer block.
2. Check `User-agent: Google-Extended` and `Disallow: /` (Sep 2023). Keep the training preference if you meant it; do not use it to diagnose Search indexing.
3. **Check blocks for OAI-SearchBot and Claude-SearchBot (2024).** Remove them if you want access for the search roles those bots serve.
4. **Check blocks for PerplexityBot (2024).** Remove them if you want inclusion in Perplexity search results.
5. Check `User-agent: *` with `Disallow: /` on production. A staging rule carried over to the live site can override your careful bot-by-bot plan.
6. **Check Cloudflare’s AI crawler control, including the default on a new domain (Jul 2024–Jul 2025).** A network block can make every robots.txt `Allow` beside the point.
7. **Check WAF rules and relevant bot IP information.** Do not assume a named user agent passes the firewall; inspect the request that was accepted or denied.
8. Check robots.txt on each relevant subdomain. A rule on the root domain does not stand in for a file on `help.`, `blog.` or `app.`.
9. Check the fetched page, not just its URL. JavaScript-only key content, consent interstitials and challenge pages returning 200 can leave a bot with little to use.
10. Check 401, 403 and 429 responses, plus soft-404s and noindex or canonical drift. A reachable-looking page may still fail at the response or extraction stage.
11. **Run a live fetch on the new or migrated domain.** The most easily missed block is the CDN default nobody remembers choosing; it can leave you chasing content changes while the answer bot never reaches the page.

Use the [45-minute curl and CDN audit](https://groas.com/post/can-chatgpt-actually-read-your-site-a-45) to confirm the failure rather than guessing from the file. If citations are the goal, keep intentional training opt-outs such as `Disallow: GPTBot`, `Disallow: ClaudeBot` and `Disallow: Google-Extended`. Remove answer-bot blocks you did not intend, adjust the CDN or WAF policy that prevents legitimate fetches, and serve the facts in a response the bot can read. The broader [technical SEO issues that hurt AI visibility](https://groas.com/post/fix-technical-seo-issues-hurt-ai-visibility) still matter, but a content rewrite cannot outrank a blocked request.

That is the order I want when groas works on earned search: find the setting, understand why it was added, change only what conflicts with the current goal, and verify the fetch before writing another page. The line from 2023 to now is not that every old block was a mistake. It is that **yesterday’s training decision can become today’s answer-access problem when nobody checks what changed around it.**

## Frequently Asked Questions

### Does blocking GPTBot keep my site out of ChatGPT answers?

No. GPTBot is OpenAI's training crawler, so a Disallow for GPTBot alone does not block ChatGPT search answers. Answer-related access uses other bots such as OAI-SearchBot. Keep the GPTBot rule if you still want the training opt-out, then keep auditing.

### If I block Google-Extended, will my pages disappear from Google Search?

No. Google-Extended is a robots.txt control for Bard and Vertex AI training, not a crawler and not a Google Search indexing control. Do not remove the training preference just because a page is absent from AI Overviews, and do not expect removing it to fix crawler or rendering problems elsewhere.

### Which OpenAI and other AI bots should I allow so my site can be cited in AI answers?

For ChatGPT search inclusion, OAI-SearchBot matters, because it serves the ChatGPT search index; ChatGPT-User handles on-demand fetches and GPTBot is the training crawler. Anthropic similarly separates ClaudeBot (training), Claude-SearchBot (search) and Claude-User, and Perplexity separates PerplexityBot (search results) from Perplexity-User. Judge each rule by the bot's job, not by the word AI.

### Can Cloudflare block AI bots even if my robots.txt allows them?

Yes. A robots.txt Allow cannot undo a network-layer block, so the CDN can turn a bot away before it reaches the page even when the file looks fine. Cloudflare's one-click Block AI Scrapers and Crawlers control from July 2024 lives under Security > Bots, and since July 1, 2025 blocking AI crawlers is the default for new domains unless owners opt in.

### Does publishing an llms.txt file make AI models access my site?

No. An /llms.txt file, a Markdown summary and links proposed by Jeremy Howard and Answer.AI in September 2024, does not override robots.txt or a firewall. Make sure its links point to the pages you want found, then test whether the relevant bots can actually fetch those pages. Its absence is not, by itself, an explanation for invisibility.

### Why is my robots.txt fine but AI bots still cannot fetch my pages?

Because the access decision may be made at another layer. A robots.txt Allow is not a firewall pass, and a user-agent string on its own does not establish who sent a request. Check the layer making the decision: inspect the WAF result, compare requests with published bot IP information where available, and test the page itself with a live fetch.

### How do I audit why my site is missing from AI answers?

Work from policy to live response: date each setting, identify what it was meant to control, check every layer that can deny the request (robots.txt on each subdomain, the CDN, the WAF), then run a live fetch to confirm what the bot receives. A fetch that never completes is enough to remove a page from an answer path, so a content rewrite cannot outrank a blocked request.

## Related Posts

- [![A foggy crystal ball beside a clipboard of green-ticked tasks: the agency sells the verifiable checklist, not a prediction of what AI will say.](https://groas.com/media/blog/36d26400cbf355ec679093c89d21061bee66c8a853abb9a2bf29d3f6433cdb55.png)](https://groas.com/post/the-aeo-retainer-swipe-file-copy-ready-t)
  ### [The AEO Retainer Swipe File: Tiers, Scope Clauses and Client Scripts](https://groas.com/post/the-aeo-retainer-swipe-file-copy-ready-t)
  [October 6, 2026 • 10 min read](https://groas.com/post/the-aeo-retainer-swipe-file-copy-ready-t)

  [Written by Alexander Perelman](https://groas.com/post/the-aeo-retainer-swipe-file-copy-ready-t)
- [![A speech-bubble-shaped bookshelf where gold 'Sponsored' gift boxes squeeze out the books, leaving two leaning sources and one fallen to the floor.](https://groas.com/media/blog/de47ce6e28adc50d217461e0f1214122bab80aded11ff6108eadf834719b9b8f.png)](https://groas.com/post/ads-are-moving-into-ai-overviews-five-pr)
  ### [AI Overview Ads Are Crowding the Answer: Five Citation Predictions for 2027](https://groas.com/post/ads-are-moving-into-ai-overviews-five-pr)
  [October 6, 2026 • 10 min read](https://groas.com/post/ads-are-moving-into-ai-overviews-five-pr)

  [Written by David](https://groas.com/post/ads-are-moving-into-ai-overviews-five-pr)
- [![Cartoon of a businessman in a sinking rowboat proudly admiring a gold-framed chart of the rising water while his bailing bucket sits unused.](https://groas.com/media/blog/267ab77791d4bf1bfac9d34e77020fcd7536c01d8309aa711f3ad08d6e7c8c1b.png)](https://groas.com/post/introducing-citeboost-platinum-watch-you)
  ### [Introducing CiteBoost Platinum: Beautiful Charts of Every AI Citation You Didn’t Earn](https://groas.com/post/introducing-citeboost-platinum-watch-you)
  [October 6, 2026 • 6 min read](https://groas.com/post/introducing-citeboost-platinum-watch-you)

  [Written by David](https://groas.com/post/introducing-citeboost-platinum-watch-you)

## The Machines Already Run Search, You Should Probably Own One

[get started free](https://groas.typeform.com/to/xC1bQNUT)

## Structured data

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://graph.groas.com/entity/groas","name":"groas","url":"https://groas.com/","logo":{"@type":"ImageObject","url":"https://groas.com/icon-512.png","width":512,"height":512},"sameAs":["https://www.linkedin.com/company/groas/"]},{"@type":"WebSite","@id":"https://groas.com/#website","name":"groas","url":"https://groas.com/","inLanguage":"en-US","publisher":{"@id":"https://graph.groas.com/entity/groas"}},{"@type":"WebPage","@id":"https://groas.com/post/from-gptbot-to-default-blocking-a-dated#webpage","url":"https://groas.com/post/from-gptbot-to-default-blocking-a-dated","name":"The AI Visibility Blocks Your Site Forgot: A Timeline from GPTBot to CDN Defaults","description":"Your site may be missing from AI answers because someone made a perfectly sensible decision in 2023 and nobody checked it again.","inLanguage":"en-US","isPartOf":{"@id":"https://groas.com/#website"},"about":{"@id":"https://graph.groas.com/entity/groas"},"primaryImageOfPage":{"@type":"ImageObject","url":"https://groas.com/media/blog/8e8379bcc6e1091943afb11a15fa233040f034f9b4ceb47cd11526669070d3cd.png"},"breadcrumb":{"@id":"https://groas.com/post/from-gptbot-to-default-blocking-a-dated#breadcrumb"},"dateModified":"2026-10-04T05:43:26.253Z","mainEntity":{"@id":"https://groas.com/post/from-gptbot-to-default-blocking-a-dated#article"}},{"@type":"BreadcrumbList","@id":"https://groas.com/post/from-gptbot-to-default-blocking-a-dated#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://groas.com/"},{"@type":"ListItem","position":2,"name":"Blog","item":"https://groas.com/blog"},{"@type":"ListItem","position":3,"name":"Google Ads News","item":"https://groas.com/category/google-ads-news"},{"@type":"ListItem","position":4,"name":"The AI Visibility Blocks Your Site Forgot: A Timeline from GPTBot to CDN Defaults","item":"https://groas.com/post/from-gptbot-to-default-blocking-a-dated"}]},{"@type":"BlogPosting","@id":"https://groas.com/post/from-gptbot-to-default-blocking-a-dated#article","url":"https://groas.com/post/from-gptbot-to-default-blocking-a-dated","mainEntityOfPage":{"@id":"https://groas.com/post/from-gptbot-to-default-blocking-a-dated#webpage"},"isPartOf":{"@id":"https://groas.com/blog#blog"},"headline":"The AI Visibility Blocks Your Site Forgot: A Timeline from GPTBot to CDN Defaults","description":"Your site may be missing from AI answers because someone made a perfectly sensible decision in 2023 and nobody checked it again. The problem is not always a new technical SEO bug. Often, it is a fossil: a robots.txt rule, CDN toggle or b…","datePublished":"2026-10-04T05:43:25.149Z","dateModified":"2026-10-04T05:43:26.253Z","image":{"@type":"ImageObject","url":"https://groas.com/media/blog/8e8379bcc6e1091943afb11a15fa233040f034f9b4ceb47cd11526669070d3cd.png"},"author":{"@id":"https://groas.com/author/david#person"},"publisher":{"@id":"https://graph.groas.com/entity/groas"},"articleSection":"Google Ads News","keywords":"Google Ads, Google Ads News","wordCount":1974,"inLanguage":"en-US"},{"@type":"Person","@id":"https://groas.com/author/david#person","name":"David","jobTitle":"Founder & CEO @ groas","image":"https://groas.com/media/blog/363be9ad5255816654dac065d01bf1c6063110b400bb38cc199570267232027c.jpg","worksFor":{"@id":"https://graph.groas.com/entity/groas"}},{"@type":"FAQPage","@id":"https://groas.com/post/from-gptbot-to-default-blocking-a-dated#faq","mainEntity":[{"@type":"Question","name":"Does blocking GPTBot keep my site out of ChatGPT answers?","acceptedAnswer":{"@type":"Answer","text":"No. GPTBot is OpenAI's training crawler, so a Disallow for GPTBot alone does not block ChatGPT search answers. Answer-related access uses other bots such as OAI-SearchBot. Keep the GPTBot rule if you still want the training opt-out, then keep auditing."}},{"@type":"Question","name":"If I block Google-Extended, will my pages disappear from Google Search?","acceptedAnswer":{"@type":"Answer","text":"No. Google-Extended is a robots.txt control for Bard and Vertex AI training, not a crawler and not a Google Search indexing control. Do not remove the training preference just because a page is absent from AI Overviews, and do not expect removing it to fix crawler or rendering problems elsewhere."}},{"@type":"Question","name":"Which OpenAI and other AI bots should I allow so my site can be cited in AI answers?","acceptedAnswer":{"@type":"Answer","text":"For ChatGPT search inclusion, OAI-SearchBot matters, because it serves the ChatGPT search index; ChatGPT-User handles on-demand fetches and GPTBot is the training crawler. Anthropic similarly separates ClaudeBot (training), Claude-SearchBot (search) and Claude-User, and Perplexity separates PerplexityBot (search results) from Perplexity-User. Judge each rule by the bot's job, not by the word AI."}},{"@type":"Question","name":"Can Cloudflare block AI bots even if my robots.txt allows them?","acceptedAnswer":{"@type":"Answer","text":"Yes. A robots.txt Allow cannot undo a network-layer block, so the CDN can turn a bot away before it reaches the page even when the file looks fine. Cloudflare's one-click Block AI Scrapers and Crawlers control from July 2024 lives under Security > Bots, and since July 1, 2025 blocking AI crawlers is the default for new domains unless owners opt in."}},{"@type":"Question","name":"Does publishing an llms.txt file make AI models access my site?","acceptedAnswer":{"@type":"Answer","text":"No. An /llms.txt file, a Markdown summary and links proposed by Jeremy Howard and Answer.AI in September 2024, does not override robots.txt or a firewall. Make sure its links point to the pages you want found, then test whether the relevant bots can actually fetch those pages. Its absence is not, by itself, an explanation for invisibility."}},{"@type":"Question","name":"Why is my robots.txt fine but AI bots still cannot fetch my pages?","acceptedAnswer":{"@type":"Answer","text":"Because the access decision may be made at another layer. A robots.txt Allow is not a firewall pass, and a user-agent string on its own does not establish who sent a request. Check the layer making the decision: inspect the WAF result, compare requests with published bot IP information where available, and test the page itself with a live fetch."}},{"@type":"Question","name":"How do I audit why my site is missing from AI answers?","acceptedAnswer":{"@type":"Answer","text":"Work from policy to live response: date each setting, identify what it was meant to control, check every layer that can deny the request (robots.txt on each subdomain, the CDN, the WAF), then run a live fetch to confirm what the bot receives. A fetch that never completes is enough to remove a page from an answer path, so a content rewrite cannot outrank a blocked request."}}]}]}
```
