

You can write the clearest answer on the internet and still give ChatGPT nothing to cite. Before you pay for another page, check whether AI crawlers can fetch the one you already have.
I used to tell clients the answer was better copy and more FAQs. I was wrong about the order. A page blocked by robots.txt or a firewall, or delivered as an empty JavaScript shell, cannot help a crawler that never sees its answer. This guide starts at that gate, follows the fetch through your edge logs, and leaves content gaps until there is evidence the plumbing works.
Most AI-visibility advice starts at step five. It hands you a content score, tells you to add clear answers and schema, and sends you off to publish. That is backwards. I have watched businesses spend three months publishing answer pages while an inherited robots.txt rule or a firewall rule kept relevant crawlers out.
Run the checks in this order:
Skip ahead and you get what many tools sell you: a long list of content gaps for pages a crawler cannot read. Start with access.
Open your robots.txt and search for GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Claude-SearchBot, PerplexityBot and Perplexity-User. Check the rules that apply to each one and to the pages you want surfaced. A Disallow: / is worth investigating, but a training-bot block is not automatically a search-bot block. The fix is precise: allow the search and live-answer access you want; keep blocks you actually mean to keep.
The part most checklists get wrong is treating every bot as one switch. OpenAI lets you allow OAI-SearchBot for ChatGPT search answers while disallowing GPTBot for training; its controls work independently, and changes can take about 24 hours to take effect. Anthropic runs separate bots with different consequences if blocked: blocking Claude-SearchBot reduces search visibility, while blocking Claude-User reduces live fetches when a user asks about a page. Perplexity likewise distinguishes PerplexityBot, which helps build its index, from Perplexity-User, which fetches during an answer. Check the names separately. Do not assume a training decision settles your citation strategy.
Then check the gate in front of robots.txt. I learned this on a home services account where GPTBot showed zero hits for a month despite an open robots.txt. Cloudflare Bot Fight Mode was bouncing AI fetchers at the edge before they touched the site. Perplexity publishes separate Cloudflare and AWS WAF rules for its bots, because a firewall can stop a fetch even when your robots rules allow it. Search WAF events for blocked requests carrying the relevant bot names, then verify identity before allow-listing them. Do not turn protection off just to see whether a crawler gets through.
Allow-list with the published bot information, including IPs and user agents where available. Anthropic warns that crude IP-blocking can stop its crawler from even reading robots.txt. That is the practical test here: if you run Cloudflare, Akamai or AWS WAF, inspect blocked requests before you change a word of content.
Getting fetched is not the same as getting read. Say you are spending $20k/month on content and your pricing page loads its numbers through a client-side API call. A browser shows the prices after JavaScript runs. A crawler that only reads the initial HTML gets a shell. In a Vercel and MERJ test, none of the major AI crawlers rendered JavaScript: ChatGPT fetched JS files in 11.5% of requests and Claude in 23.84%, but executed none of it. The script holds the answer, so the crawler leaves without the answer.
That creates partial invisibility, which is harder to spot than a blocked page. Your page may look complete in a browser while pricing, specs, comparison grids, reviews, FAQs or location details are absent from its raw HTML. Those are often the details a buyer asks for. In one controlled test across 1,062 pages, GPTBot crawled 748 HTML pages and zero JS-injected pages; after JS links became HTML links, it reached 250 new pages in under three hours.
Do the unglamorous check: view source rather than Inspect, disable JavaScript, and curl the raw HTML. Look for the actual price, spec or answer, not merely its heading or a placeholder where a widget should load. If the useful text is missing, fix delivery before you rewrite the prose.
Even server-rendered pages can lose context around the edges. I have seen robots.txt files that allow the page but disallow /assets/, /api/ or a CDN folder holding product data. The bot reaches the page, then cannot fetch a resource the answer depends on. Check the paths that deliver visible text and product data, along with images whose alt text supplies useful context. Keep intentional blocks on carts, admin areas, account pages and internal search results. A sitemap tells crawlers where to look; it does not give them permission to read.

Do not wait for a framework debate to fix a money page. If the price, spec table or comparison appears only after a script runs, put the answer in the initial HTML. Keep the interactive widget if it helps visitors. Just make sure a raw HTML fetch already contains the same substantive answer they see on screen.
Anyone can set a user agent to GPTBot. I have seen dashboards count hundreds of apparent GPTBot hits that came from scrapers, SEO tools and one very persistent competitor. Count those as proof of AI access and you can miss the fact that the real fetcher never arrived. Read verified fetches, not bot-shaped names in a chart.
For Anthropic, check the source IP against its published list instead of trusting the user agent. Perplexity publishes IP JSON for both bots so you can identify them at the WAF. Match the bot, IP and requested URL in your edge logs. Then look at the response: a verified request that received a block or an empty page is not a successful read.
This is why the check belongs at the edge. The firewall sees requests that never reach your application logs. Build one filtered view of verified AI bot requests to real pages, with their responses, and set unverified claims aside. If that view is empty, investigate access before buying a content tool.

Now the bot names become useful. Training crawlers such as GPTBot and ClaudeBot gather material for training; their visits are not evidence that a page will be cited in a current answer. Search bots such as OAI-SearchBot and Claude-SearchBot matter to search-grounded visibility. Live fetchers include ChatGPT-User, which can fetch a page because someone asked ChatGPT to visit it, and Perplexity-User, which can fetch during an answer. These are different kinds of activity. Do not roll them into one impressive-looking traffic total.
Pull 30 days of edge logs and group verified requests by bot and URL. Separate search-bot visits from live fetches. If OAI-SearchBot repeatedly reaches your blog but never your pricing or comparison pages, inspect access, links and raw HTML on those money pages before drafting another opinion piece. If Perplexity-User repeatedly fetches a comparison page, make that page an early candidate for review: the answer needs to be accurate, visible and easy to quote.
A live fetch does not tell you the user's exact question, prove they are a buyer or prove your page was cited. It does tell you that the URL was requested in a more immediate context than a training crawl. Use that signal to decide which existing pages to inspect first, then check their answers yourself.
Only now have you earned the right to plan new content. Review your recent live-fetched URLs, then separately collect the buyer questions you need to answer: price versus a competitor, best fit for a use case, pros and cons, time to get started, and whether an integration works. Your logs give you URLs, not a transcript of the prompts behind them. Do not pretend otherwise.
Map each buying question to the best existing URL. You may find three posts answering an easy introductory question and no page that addresses the buying decision. If one page is carrying several distinct decisions, check whether a reader can find a direct answer to each. If no page answers a question, that is a genuine gap. Prioritise it against the pages already receiving relevant fetches rather than treating every missing topic as equally urgent.
The fix is not automatically more words. Put the decision and its answer in plain HTML near the top of the page; include a number when a number settles the point, and write a comparison a model can quote without guessing. That is the useful lesson in why AI assistants cite some businesses and skip others: give them a passage that actually answers the question, not a homepage that gestures at it. Once access and rendering are clean, use the passage-first approach to getting cited to improve the existing pages before commissioning 40 new posts.
I would ask a tool to make the access problem visible and help confirm the repair. Can it show verified fetches by bot and page from edge logs? Can it identify the robots.txt rule, firewall event or missing raw-HTML answer behind a failure? Can you check that a fresh request reaches the page and receives useful content afterward? That is work you can act on.
Rankings and citation counts are worth monitoring, but they are not substitutes for those checks. If a platform only charts mentions while your pricing page serves a blank shell to ClaudeBot, skip it. Start with the steps for fixing technical issues that hurt AI visibility and judge the work by what changes: verified access, readable answers and, ultimately, the business outcomes you care about. Charts are cheap. Execution is the product.

Today, do the 15-minute version: open /robots.txt, open your WAF blocked-requests log, and curl one pricing page to inspect its raw HTML. If a relevant bot is blocked or the answer is missing, you have your first fix. Do that before you write another page. A citation is not guaranteed by a fetch, but a crawler cannot cite an answer it never gets to read.
In welke volgorde moet je een AI-zichtbaarheidsaudit uitvoeren?
Begin met toegang: kunnen de AI-crawlers je pagina's überhaupt bereiken? Daarna pas kijken naar wat de bot op de pagina ziet, of de fetches van echte bots komen, welke pagina's live-fetches krijgen, en pas als laatste welke koopvragen nog een pagina nodig hebben. Een contentplan schrijven terwijl een robots.txt- of firewallregel de crawler blokkeert, verspilt maanden werk.
Blokkeer ik ChatGPT als ik GPTBot in mijn robots.txt blokkeer?
Niet automatisch. OpenAI laat je OAI-SearchBot toestaan voor ChatGPT-zoekantwoorden terwijl je GPTBot blokkeert voor training; de instellingen werken onafhankelijk en wijzigingen kunnen ongeveer 24 uur duren. Anthropic en Perplexity hebben vergelijkbare aparte bots, dus controleer elke botnaam los en neem geen trainingsbeslissing als citatiebeleid.
Kunnen AI-crawlers geblokkeerd worden terwijl mijn robots.txt ze toestaat?
Ja. Een firewall of WAF kan een fetch tegenhouden voordat de site of robots.txt überhaupt bereikt wordt; het artikel noemt een account waar Cloudflare Bot Fight Mode GPTBot een maand lang tegenhield. Zoek in je WAF-gebeurtenissen naar geblokkeerde verzoeken met de betreffende botnamen en verifieer de identiteit voordat je ze toestaat.
Kunnen AI-crawlers JavaScript uitvoeren op mijn website?
Nee. In een test van Vercel en MERJ heeft geen enkele grote AI-crawler JavaScript uitgevoerd: ChatGPT haalde in 11,5% van de verzoeken JS-bestanden op en Claude in 23,84%, maar voerde geen van beide uit. Staat je prijs, spec of antwoord alleen na het uitvoeren van een script, zet het dan in de initiële HTML.
Hoe controleer ik of een AI-crawler mijn pagina-inhoud echt kan lezen?
Bekijk de broncode in plaats van Inspect, schakel JavaScript uit en curl de rauwe HTML. Zoek naar de echte prijs, specificatie of het antwoord, niet naar een kop of placeholder waar een widget moet laden. Controleer ook of paden als /api/ of een CDN-map niet geblokkeerd worden terwijl de pagina het antwoord daaruit haalt.
Hoe weet ik of een GPTBot-bezoek in mijn logs echt van OpenAI komt?
Vertrouw niet op de user agent, want iedereen kan die instellen. Vergelijk het bron-IP met de gepubliceerde IP-lijsten van de aanbieder, bijvoorbeeld de IP-JSON die Perplexity publiceert, en koppel bot, IP en opgevraagde URL in je edge-logs. Een geverifieerd verzoek dat een blokkade of lege pagina krijgt, is bovendien geen succesvolle lezing.
Betekent een bezoek van GPTBot dat ChatGPT mijn pagina zal citeren?
Nee. GPTBot en ClaudeBot verzamelen materiaal voor training; dat zegt niets over citatie in een actueel antwoord. Zoekbots zoals OAI-SearchBot en Claude-SearchBot zijn relevanter voor zoekzichtbaarheid, en live fetchers zoals ChatGPT-User en Perplexity-User halen een pagina op tijdens een antwoord. Groepeer geverifieerde verzoeken per bot en URL in plaats van alles in één totaal te gooien.