---
title: "The AI Crawler Swipe File: Rules and Page Blocks for Live-Answer Citations"
description: "Dedicated strategists run your account and a proprietary engine trained on $500B in profitable ad spend optimises the execution underneath them"
image: "https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6ac1e56b51be33048b4b224b_729c8230-e3cd-45ca-8150-54c01d182578.png"
---

October 4, 2026

•

11

min read

# The AI Crawler Swipe File: Rules and Page Blocks for Live-Answer Citations

![Young man with curly hair wearing a black shirt outdoors against green foliage background.](https://cdn.prod.website-files.com/6821efca072e48f6f495a47e/68562d390107b3921a6e3d68_1743932904108.jpg)

**Alexander Perleman**, Head Of Product @ groas
Ex-Goldman Sachs and Stanford Computer Science

**Email: alex@groas.com**

[**LinkedIn: https://www.linkedin.com/in/alexander-433793253/**](https://www.linkedin.com/in/alexander-433793253/)

![Cover image for: The AI Crawler Swipe File: Rules and Page Blocks for Live-Answer Citations](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6ac1e56b51be33048b4b224b_729c8230-e3cd-45ca-8150-54c01d182578.png)

Most sites that “block AI” are blocking the wrong bot. I spent a year disallowing GPTBot and thinking I had opted out of ChatGPT, while the user agent that could fetch a page for a buyer’s live question was a different one entirely.

#### Bot map: know which door each rule closes

**When to use this:** Before changing `/robots.txt`. A training crawler collects content that may feed a model; a search bot indexes pages that may surface in answers; a user fetcher requests a page because someone asked about it. Those are different jobs, so one blanket rule is a bad policy.

| Provider | Training control | Search or indexing | Live user fetch |
| --- | --- | --- | --- |
| OpenAI | `GPTBot` | `OAI-SearchBot` | `ChatGPT-User` |
| Anthropic | `ClaudeBot` | `Claude-SearchBot` | `Claude-User` |
| Perplexity | — | `PerplexityBot` | `Perplexity-User` |
| Google | `Google-Extended` controls the use described in Google’s documentation | Regular Googlebot feeds Search and AI Overviews | — |

[OpenAI documents its three user agents separately](https://developers.openai.com/api/docs/bots): blocking GPTBot does not also block OAI-SearchBot. ChatGPT-User is a user-initiated fetch, not an automatic crawler, and robots.txt rules may not apply to it. [Perplexity likewise separates PerplexityBot from Perplexity-User](https://docs.perplexity.ai/guides/bots); its user-initiated fetch generally ignores robots.txt. [Anthropic lists ClaudeBot, Claude-SearchBot and Claude-User separately](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler), with all three honoring robots.txt. [Google-Extended does not control inclusion or ranking in Google Search or AI Overviews](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers).

**Adjustment that matters:** A line saying `Allow: /` is not a citation switch. It states your crawling policy for an agent that honors it. Keep the indexing bots and live fetchers in view when you choose that policy.

#### robots.txt swipe A: leave the citation path open

**When to use this:** You want maximum eligibility and have no policy against the training use these controls cover. **Affects:** The listed training crawlers, search bots and user fetchers, subject to each provider’s handling of robots.txt. Use this as your policy, not as a few lines pasted beneath an existing site-wide block.

```txt
# Variant A: open to the listed agents
User-agent: GPTBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /

User-agent: Google-Extended
Allow: /
```

**Adjustment that matters:** Review paths you do not want crawled, such as admin, cart and internal-search pages, without blocking the service or product pages you want found. I used to reach for broad wildcard rules when I meant to exclude junk URLs. That is a clumsy way to protect a handful of pages, and it can shut the door on the pages doing the selling.

#### robots.txt swipe C: the two-line mistake

**When to use this:** As a warning, not a template, if someone says the site has “blocked AI.” **Affects:** Automatic crawlers broadly, including Googlebot and search bots you may need for visibility. **Do not use it if you want those crawlers to reach the site.**

```txt
# Variant C: accidental site-wide block
User-agent: *
Disallow: /
```

**Adjustment that matters:** Fetch your live `/robots.txt` and check for this before adding a carefully written agent-specific rule. I have seen a blanket block sit under a comment that says `# block AI bots`. The comment sounds selective. The rule is not. A wildcard `Disallow: /` is a far bigger decision than opting out of training.

Google has its own version of this confusion. If you disallow `Google-Extended`, you have not opted out of AI Overviews: [that control does not determine Search or AI Overviews inclusion](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers). Say you are spending $20k a month on content intended to win local-service answers. Blocking Google-Extended and then seeing your pages in Overviews is not a failed deployment. Blocking Googlebot to leave Search would be a different, much costlier tradeoff. **Do not mistake a training-use control for a Search opt-out.**

![Cartoon of a training crawler stopped at a door while a live-answer fetcher enters](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6ac1e56c51be33048b4b22a2_4f24508c-ab99-4834-af39-478f09297ce1.png)

#### Schema swipe: identify the business, then check the page

**When to use this:** On a business homepage. **Affects:** Structured-data parsers and entity identification, including Google’s understanding of the business. JSON-LD labels facts on the page; it does not force an answer engine to quote them. Put the completed block in a JSON-LD script in `<head>` or add it through your SEO plugin. Replace every bracketed placeholder before publishing.

```json
{
  "@context": "https://schema.org",
  "@type": "Organization",
  "name": "[BUSINESS NAME]",
  "url": "https://[YOUR-DOMAIN]/",
  "logo": "https://[YOUR-DOMAIN]/logo.png",
  "sameAs": [
    "https://www.linkedin.com/company/[HANDLE]/",
    "https://www.crunchbase.com/organization/[HANDLE]",
    "https://[REVIEW-SITE]/[YOUR-PROFILE]"
  ]
}
```

**Adjustment that matters:** Keep only profiles you control and keep current. Delete a `sameAs` entry rather than leaving a dead or stale profile in the array. Three useful links beat ten that send a parser looking for the wrong business.

#### Schema swipe: put the offer beside the visible price

**When to use this:** On a core product or service page. **Affects:** Price and offer extraction. Choose either `Product` or `Service` for `@type`, replace the bracketed values, and make the published page agree with them. The code is a starting block, not a claim that every product has a rating.

```json
{
  "@context": "https://schema.org",
  "@type": "[Product or Service]",
  "name": "[PRODUCT NAME]",
  "description": "[One sentence: who it is for and what it does.]",
  "brand": { "@type": "Brand", "name": "[BUSINESS NAME]" },
  "offers": {
    "@type": "Offer",
    "price": "[99.00]",
    "priceCurrency": "[USD]",
    "availability": "https://schema.org/InStock",
    "url": "https://[YOUR-DOMAIN]/[PRODUCT-SLUG]/"
  },
  "aggregateRating": {
    "@type": "AggregateRating",
    "ratingValue": "[4.8]",
    "reviewCount": "[214]"
  }
}
```

![Photograph of labelled moving boxes with facts on the outside for AI parsers](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6ac1e56c51be33048b4b22a5_7233476a-ee89-4c63-8e2e-132e847f6820.png)

**Adjustment that matters:** If the page says “starting at $149” and the block says `$99`, fix the disagreement rather than hoping a parser picks the nicer number. I remove `aggregateRating` unless the reviews live on that URL. Do not paste the sample rating into a real page because it happens to look convincing.

#### Schema swipe: mark up answers buyers can already read

**When to use this:** On a pricing, process or comparison page with a visible question and answer. **Affects:** Structured interpretation of that page, not a promise of a Google FAQ feature. [FAQPage remains documented vocabulary](https://developers.google.com/search/docs/appearance/structured-data/faqpage); use it because the answer is useful, not because you expect a special search result.

```json
{
  "@context": "https://schema.org",
  "@type": "FAQPage",
  "mainEntity": [{
    "@type": "Question",
    "name": "[Question buyers actually ask?]",
    "acceptedAnswer": {
      "@type": "Answer",
      "text": "[40–60 word direct answer with one number.]"
    }
  }]
}
```

**Adjustment that matters:** Match the question and answer to visible copy word for word. If you change the page, change the markup. Otherwise you have built a tidy label for an answer the reader cannot find.

![Close-up of an answer-first paragraph highlighted on a laptop screen](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6ac1e56b51be33048b4b2293_672c1cd0-0bdd-47ab-b671-e3ac19781ee2.png)

#### Content swipe: give an answer engine a sentence worth lifting

**When to use this:** At the top of a service, product or pricing page you want quoted. **Affects:** The readable page content that live fetchers and search systems may use. In the [Princeton GEO benchmark](https://arxiv.org/abs/2311.09735), citations, statistics and quotations improved visibility in the reported tests, while keyword stuffing hurt it. The useful lesson is not to sprinkle numbers on empty claims. It is to make the answer, its evidence and its limits easy to read.

```md
[BUSINESS NAME] [does X] for [SPECIFIC BUYER] in [LOCATION/SCOPE].
Clients pay [PRICE RANGE] and typically see [ONE MEASURED OUTCOME] within [TIMEFRAME], according to [NAMED SOURCE].
Best for [USE CASE]. Not for [EXCLUDED USE CASE].
```

**Adjustment that matters:** Keep the measured outcome only if you have one you can support; otherwise rewrite that sentence around what you can state plainly. Put the price and timeframe near the answer instead of opening with company history. A buyer asking what you do should not need to scroll past “founded in” to find out.

#### Content swipe: make the comparison easy to inspect

**When to use this:** On pricing, alternatives and versus pages where buyers ask which option fits. **Affects:** The comparison a reader or answer engine can extract from the page. Keep the rows grounded in your actual offer, and show why it is the stronger fit for the buyer you serve.

```md
| Option | Best for | Price | Tradeoff |
|---|---|---|---|
| [YOUR PRODUCT] | [SPECIFIC BUYER + JOB] | [PRICE] | [REAL LIMIT] |
| [ALTERNATIVE 1] | [DIFFERENT BUYER + JOB] | [PRICE] | [WHY IT MISSES YOUR BUYER'S NEED] |
| [ALTERNATIVE 2] | [DIFFERENT BUYER + JOB] | [PRICE] | [WHY IT MISSES YOUR BUYER'S NEED] |
```

**Adjustment that matters:** Name your own limit without pretending a weaker alternative does your buyer’s job better. A table where your product wins every cell reads like an ad. A table that makes the decision legible is much more useful.

#### Content swipe: say who should walk away

**When to use this:** Directly under pricing or a comparison table. **Affects:** The answer a buyer gets about fit, and the enquiries your team has to sort later.

```md
**Who this is not for:** [BUSINESS NAME] is not for [SPECIFIC BAD-FIT SEGMENT] that need [REASON].
Those buyers should [MORE SUITABLE NEXT STEP]. If you are [GOOD-FIT SIGNAL], [NEXT STEP].
```

**Adjustment that matters:** Name a real exclusion, not “anyone who doesn’t value quality.” Budget, timeline or location will do more work. “Not for single-location shops under $3k/month that need same-day install” tells a buyer something; vague selectivity just poses for the camera.

#### llms.txt swipe: offer a map, not a magic switch

**When to use this:** When long docs or product pages would benefit from a clean list of destinations. **Affects:** Helpers that choose to read the file; it does not control training, replace robots.txt or repair an unreadable page. [The *llms.txt* proposal](https://llmstxt.org/) describes a Markdown file at `/llms.txt` with a site-name heading, a short summary and lists of links. Check whether your pages can be read first using [Before You Write for ChatGPT, Check Whether AI Can Read Your Site](https://groas.com/post/before-you-write-another-page-for-chatgp).

```md
# [BUSINESS NAME]

> [One sentence: what you sell, for whom, in one region or scope.]

## Core pages
- [Pricing](https://[YOUR-DOMAIN]/pricing/): [plans and price ranges in one line]
- [Services](https://[YOUR-DOMAIN]/services/): [who each service fits]
- [Comparisons](https://[YOUR-DOMAIN]/vs/): [how you differ from alternatives]
- [FAQ](https://[YOUR-DOMAIN]/faq/): [delivery, timeline, and support answers]

## Optional
- [About](https://[YOUR-DOMAIN]/about/): [founded, team size, location]
```

![Minimalist diagram of a small llms.txt map beside a large robots.txt gate](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6ac1e56b51be33048b4b2290_642253ea-5695-4af3-8bf2-e5fc3a589af0.png)

**Adjustment that matters:** Link to pages that answer the question promised by each label. If the pricing page hides pricing, a neat map only gets a reader to the disappointment faster.

#### Verification swipe: check the deployed file, not the draft

**When to use this:** After changing robots.txt or publishing any page block. **Affects:** Your ability to tell a deployment problem from a page-content problem. I start with logs and the live file, not with a victory lap in a dashboard.

1. Fetch `/robots.txt` fresh. Confirm the intended policy is live and no unintended wildcard `Disallow: /` remains.
2. Check server logs for `GPTBot`, `OAI-SearchBot`, `ChatGPT-User`, `ClaudeBot`, `Claude-User`, `Claude-SearchBot`, `PerplexityBot` and `Google-Extended`. I inspect a 14-day window; an absent user agent is a reason to investigate, not proof that a rule failed.
3. Validate the Organization and product or service JSON-LD. Compare each published value with what a buyer can see on the page.
4. Run the same five buyer prompts in ChatGPT, Claude and Perplexity. Record which pages get cited, if any. If pages are fetched but hard to use, work through [How to Fix Technical SEO Issues That Hurt AI Visibility](https://groas.com/post/fix-technical-seo-issues-hurt-ai-visibility).

**Adjustment that matters:** Keep the prompts and checks consistent when you repeat them. Otherwise a new answer can look like progress when you have merely asked a different question.

#### Final swipe: the training opt-out I would deploy first

**When to use this:** You object to the training use covered by these controls but still want the listed search bots and live-answer fetchers able to reach your pages. **Affects:** GPTBot, ClaudeBot and Google-Extended on one side; OAI-SearchBot, Claude-SearchBot, PerplexityBot and the listed user fetchers on the other. I use this variant on my own test sites. Replace your conflicting rules rather than appending it beneath them.

```txt
# Variant B: block listed training controls; allow search and live answers
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: Claude-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /
```

**Adjustment that matters:** Do not put a `Crawl-delay` on an allowed Claude fetcher when you want a live answer to reach the page promptly; Anthropic honors that directive. **This is the artefact I rate most useful** because it turns “block AI” into an actual policy: training controls are disallowed, while the search and live-answer paths remain open where robots.txt governs them. It cannot promise a citation. It can stop you from closing the wrong door. For the next page blocks to build, use [Copy-Paste AEO: Five Page Blocks Built for AI Answers](https://groas.com/post/copy-paste-aeo-the-page-structures-schem).

#### Frequently asked questions

**If I block GPTBot, does that also block ChatGPT from citing my site?**

No. GPTBot is OpenAI's training crawler, while OAI-SearchBot handles indexing and ChatGPT-User performs user-initiated fetches. Blocking GPTBot does not also block OAI-SearchBot, and ChatGPT-User is not an automatic crawler, so robots.txt rules may not apply to it.

**Does blocking Google-Extended opt my site out of AI Overviews?**

No. Google-Extended only controls the training-related use described in Google's documentation and does not determine inclusion or ranking in Google Search or AI Overviews. Regular Googlebot feeds both Search and AI Overviews, so disallowing Google-Extended is not a Search opt-out.

**Is a wildcard Disallow: / in robots.txt just an AI opt-out?**

No, a wildcard User-agent: \* with Disallow: / blocks all crawlers, including Googlebot and the search bots a site needs for visibility. Comments like "block AI bots" can sound selective, but the rule applies to every agent that honors robots.txt.

**How can I block AI training but still let search bots and live answer fetchers reach my site?**

Use agent-specific rules that disallow GPTBot, ClaudeBot and Google-Extended, while allowing OAI-SearchBot, ChatGPT-User, Claude-User, Claude-SearchBot, PerplexityBot and Perplexity-User. Also avoid putting a Crawl-delay on an allowed Claude fetcher, because Anthropic honors that directive.

**Can I paste a sample aggregateRating into my product schema even if the page has no reviews?**

No. Keep the schema aligned with what the published page actually shows: if the page says starting at $149 but the block says $99, fix the mismatch rather than hoping a parser picks the nicer number. Remove aggregateRating unless the reviews live on that URL.

**Should FAQPage schema say something different from the visible page content?**

No. Match the question and answer in the markup to the visible copy word for word, and update the markup whenever the page changes. Otherwise the markup labels an answer readers cannot find on the page.

**What kind of content gets cited by answer engines?**

Content that opens with a direct answer and makes its evidence and limits easy to read. In the Princeton GEO benchmark, citations, statistics and quotations improved visibility in the reported tests, while keyword stuffing hurt it.

**Does an llms.txt file block AI training or replace robots.txt?**

No. The llms.txt proposal describes a Markdown file at /llms.txt with a site-name heading, a short summary and lists of links, which helpers may choose to read. It does not control training, replace robots.txt or repair an unreadable page.

## Related Posts

[![A foggy crystal ball beside a clipboard of green-ticked tasks: the agency sells the verifiable checklist, not a prediction of what AI will say.](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6ac49220d54ff28579b596c5_88e2eb67-b856-4d22-962a-a4e3dc5d7df6.png) ##### The AEO Retainer Swipe File: Tiers, Scope Clauses and Client Scripts October 6, 2026 • 11 min read Written by Alexander Perelman](https://groas.com/post/the-aeo-retainer-swipe-file-copy-ready-t)

[![Clay figure pours a wheelbarrow of articles into a wall funnel whose pipe is disconnected, so the pages pile on the floor beside an unused wrench.](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6ac48e5808a82e2ea5358099_e3125d33-6dbe-4978-b7a4-40d772753208.png) ##### Can ChatGPT Read Your Site? A 45-Minute Check Before You Write More AEO Content October 6, 2026 • 11 min read Written by Alexander Perelman](https://groas.com/post/can-chatgpt-even-read-your-site-a-45-min)

[![Cartoon of a marketer noting 'position two' after one pull of a slot machine whose reels are already spinning to a new result.](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6ac48e47904ce2271e972a6f_9b960565-71b9-49f2-906a-482b1c4c2a39.png) ##### Six PPC Habits That Misread AI Visibility October 6, 2026 • 12 min read Written by David](https://groas.com/post/i-measured-ai-visibility-like-a-ppc-acco)

![](https://cdn.prod.website-files.com/6821efca072e48f6f495a47e/6823c03839b2b618430d5ae6_Light-p-500.png)

[![White stylized owl eyes with green background icon.](https://cdn.prod.website-files.com/6821efca072e48f6f495a47e/6a68a036a5dd5627409d3579_Groas%20Icon.png)](https://groas.com/post/the-ai-crawler-swipe-file-copy-paste-rob#)

[contact](mailto:scale@groas.ai)

Explore

[Home](https://groas.com/)[Philosophy](https://groas.com/our-philosophy)[Blog](https://groas.com/blog)

What We Do

[Paid Search](https://groas.com/paid-search)[Earned Search](https://groas.com/earned-search)

Who We Serve

[Businesses](https://groas.com/for-businesses)[Agencies](https://groas.com/for-agencies)

MCP

[ChatGPT](https://chatgpt.com/plugins/plugin_asdk_app_6a9c2d7108f08191beb5b3d736be8655?q=groas)[Claude](https://claude.ai/directory/mcp-groas-ai)

Legal

[Privacy Policy](https://groas.com/legal/privacy-policy)[Terms of Service](https://groas.com/legal/terms)

Get Started

[Apply](https://groas.typeform.com/to/xC1bQNUT)

© 2026 groas 🇺🇸

[![](https://cdn.prod.website-files.com/62434fa732124a0fb112aab4/62434fa732124a389912aad8_linkedin%20small.svg)](https://www.linkedin.com/company/groas/)

```json
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@type": "BlogPosting",
      "@id": "https://www.groas.com/post/the-ai-crawler-swipe-file-copy-paste-rob#article",
      "headline": "The AI Crawler Swipe File: Rules and Page Blocks for Live-Answer Citations",
      "description": "",
      "url": "https://www.groas.com/post/the-ai-crawler-swipe-file-copy-paste-rob",
      "image": {
        "@type": "ImageObject",
        "url": "https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6ac1e56b51be33048b4b224b_729c8230-e3cd-45ca-8150-54c01d182578.png"
      },
      "thumbnailUrl": "https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6ac1e56b51be33048b4b224b_729c8230-e3cd-45ca-8150-54c01d182578.png",
      "datePublished": "2026-10-04T05:34:37.319Z",
      "dateModified": "2026-10-04T05:34:37.319Z",
      "inLanguage": "en-US",
      "isAccessibleForFree": true,
      "articleSection": "Google Ads News",
      "keywords": "Google Ads, Google Ads News",
      "mainEntityOfPage": {
        "@type": "WebPage",
        "@id": "https://www.groas.com/post/the-ai-crawler-swipe-file-copy-paste-rob"
      },
      "author": { "@id": "https://www.groas.com/author/david#person" },
      "publisher": { "@id": "https://www.groas.com#organization" },
      "about": [
        { "@type": "Thing", "name": "Google Ads" },
        { "@type": "Thing", "name": "Pay-per-click advertising" },
        { "@type": "Thing", "name": "Performance marketing" }
      ],
      "mentions": [
        { "@type": "Organization", "name": "groas", "url": "https://www.groas.com" },
        { "@type": "Thing", "name": "Google Ads" }
      ],
      "speakable": {
        "@type": "SpeakableSpecification",
        "cssSelector": ["h1"]
      }
    },
    {
      "@type": "Person",
      "@id": "https://www.groas.com/author/david#person",
      "name": "David",
      "description": "",
      "jobTitle": "Founder &amp; CEO @ groas",
      "email": "",
      "url": "https://www.groas.com/author/david",
      "image": {
        "@type": "ImageObject",
        "url": "https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/686a2a76611e759fd8d8a3fd_David%20LinkedIn%20Profile.jpeg"
      },
      "sameAs": [
        ""
      ],
      "knowsAbout": [
        "Google Ads",
        "Performance Max",
        "AI Max",
        "Pay-per-click advertising",
        "Conversion tracking",
        "Bid management",
        "Search advertising"
      ],
      "worksFor": { "@id": "https://www.groas.com#organization" }
    },
    {
      "@type": "Organization",
      "@id": "https://www.groas.com#organization",
      "name": "groas",
      "alternateName": "groas.com",
      "url": "https://www.groas.com",
      "logo": {
        "@type": "ImageObject",
        "url": "https://cdn.prod.website-files.com/6821efca072e48f6f495a47e/6834a970e03a7016560e3515_Logo%20Design%20256x256.png"
      },
      "description": "Dedicated strategists run your account and a proprietary engine trained on $500B in profitable ad spend optimises the execution underneath them",
      "sameAs": [
        "https://www.linkedin.com/company/groas"
      ],
      "knowsAbout": [
        "Google Ads",
        "Performance Max",
        "AI Max for Search",
        "Pay-per-click advertising",
        "Autonomous campaign management",
        "AI advertising agents"
      ],
      "areaServed": "Worldwide"
    },
    {
      "@type": "BreadcrumbList",
      "itemListElement": [
        {
          "@type": "ListItem",
          "position": 1,
          "name": "Home",
          "item": "https://www.groas.com"
        },
        {
          "@type": "ListItem",
          "position": 2,
          "name": "Blog",
          "item": "https://www.groas.com/blog"
        },
        {
          "@type": "ListItem",
          "position": 3,
          "name": "Google Ads News",
          "item": "https://www.groas.com/category/Google Ads News"
        },
        {
          "@type": "ListItem",
          "position": 4,
          "name": "The AI Crawler Swipe File: Rules and Page Blocks for Live-Answer Citations",
          "item": "https://www.groas.com/post/the-ai-crawler-swipe-file-copy-paste-rob"
        }
      ]
    }
  ]
}
```

```json
{"@context":"https://schema.org","@type":"FAQPage","mainEntity":[{"@type":"Question","name":"If I block GPTBot, does that also block ChatGPT from citing my site?","acceptedAnswer":{"@type":"Answer","text":"No. GPTBot is OpenAI's training crawler, while OAI-SearchBot handles indexing and ChatGPT-User performs user-initiated fetches. Blocking GPTBot does not also block OAI-SearchBot, and ChatGPT-User is not an automatic crawler, so robots.txt rules may not apply to it."}},{"@type":"Question","name":"Does blocking Google-Extended opt my site out of AI Overviews?","acceptedAnswer":{"@type":"Answer","text":"No. Google-Extended only controls the training-related use described in Google's documentation and does not determine inclusion or ranking in Google Search or AI Overviews. Regular Googlebot feeds both Search and AI Overviews, so disallowing Google-Extended is not a Search opt-out."}},{"@type":"Question","name":"Is a wildcard Disallow: / in robots.txt just an AI opt-out?","acceptedAnswer":{"@type":"Answer","text":"No, a wildcard User-agent: * with Disallow: / blocks all crawlers, including Googlebot and the search bots a site needs for visibility. Comments like \"block AI bots\" can sound selective, but the rule applies to every agent that honors robots.txt."}},{"@type":"Question","name":"How can I block AI training but still let search bots and live answer fetchers reach my site?","acceptedAnswer":{"@type":"Answer","text":"Use agent-specific rules that disallow GPTBot, ClaudeBot and Google-Extended, while allowing OAI-SearchBot, ChatGPT-User, Claude-User, Claude-SearchBot, PerplexityBot and Perplexity-User. Also avoid putting a Crawl-delay on an allowed Claude fetcher, because Anthropic honors that directive."}},{"@type":"Question","name":"Can I paste a sample aggregateRating into my product schema even if the page has no reviews?","acceptedAnswer":{"@type":"Answer","text":"No. Keep the schema aligned with what the published page actually shows: if the page says starting at $149 but the block says $99, fix the mismatch rather than hoping a parser picks the nicer number. Remove aggregateRating unless the reviews live on that URL."}},{"@type":"Question","name":"Should FAQPage schema say something different from the visible page content?","acceptedAnswer":{"@type":"Answer","text":"No. Match the question and answer in the markup to the visible copy word for word, and update the markup whenever the page changes. Otherwise the markup labels an answer readers cannot find on the page."}},{"@type":"Question","name":"What kind of content gets cited by answer engines?","acceptedAnswer":{"@type":"Answer","text":"Content that opens with a direct answer and makes its evidence and limits easy to read. In the Princeton GEO benchmark, citations, statistics and quotations improved visibility in the reported tests, while keyword stuffing hurt it."}},{"@type":"Question","name":"Does an llms.txt file block AI training or replace robots.txt?","acceptedAnswer":{"@type":"Answer","text":"No. The llms.txt proposal describes a Markdown file at /llms.txt with a site-name heading, a short summary and lists of links, which helpers may choose to read. It does not control training, replace robots.txt or repair an unreadable page."}}]}
```
