Our edge logs recorded 11,793 bot hits in 30 days. During live queries, Claude-User fetched our robots.txt file 1,160 times. Meanwhile, one guide to 2026 YouTube ad formats drew 273 live-answer fetches from OpenAI and Perplexity user bots; our category hubs drew zero. That gap is the problem with most advice about getting cited by AI assistants. It tells businesses to rank first, add schema, publish an llms.txt file, and wait. But a citation is a fetch-and-extract event, not a ranking reward. The bot needs to reach a page and find a passage it can use. A fetch does not guarantee a citation, but prestige is a poor substitute for either step.
Myth 1: “Rank #1 on Google and ChatGPT will cite you”
I understand why performance marketers cling to this one. Google rankings are familiar, measurable, and already have a line in the retainer. Surely an assistant starts with the top result and paraphrases it?
The overlap is much thinner than that story suggests. In an Ahrefs study of 15,000 commercial and informational prompts, only 12% of URLs cited by AI assistants ranked in Google’s top 10 for the exact prompt. Perplexity’s cited sources overlapped with Google’s top 10 results 28.6% of the time; ChatGPT’s overlapped just 6.1% of the time. A top-10 position may help a page get noticed elsewhere, but it is plainly not a citation queue.
The mechanism matters. An assistant can search for a narrower version of the user’s question, retrieve candidate pages, and pull out passages that answer it. A sprawling guide may rank for the head term while burying the useful answer beneath an introduction, a table of contents, and several thousand words of warm-up. A smaller reference page may put the answer where the retrieval system can use it.
That changes what I would inspect on a page that ranks but never seems to appear in answers. I would not start by rewriting its title to chase one more position. I would read the section that actually answers the question. Does it state the answer on the page, or does it send the reader through three links and a sales pitch? If the useful sentence is hard for a person to locate, a ranking report will not make it easier to extract.
Keep ranking work, but give each important answer a compact, self-contained passage. A Google position alone does not write that passage for you.
Myth 2: “Schema markup is the citation switch”
This pitch has the neatness of a technical shortcut. Wrap your FAQs and product details in JSON-LD, the thinking goes, and AI engines will recognize your authority. It also gives consultants something tidy to invoice for.
The observed citation lift is less tidy. Ahrefs tracked 1,885 pages that added JSON-LD schema against 4,000 control pages over seven months. ChatGPT citations shifted by 2.2%, Google AI Overviews dropped 4.6%, and Perplexity showed no lift from the markup.

There is a useful distinction between finding text and rewarding its wrapper. In practitioner testing by Mark Williams-Cook, ChatGPT and Perplexity extracted facts placed inside invalid JSON-LD scripts. The broken markup did not stop the text from being read. It also did not demonstrate that valid schema earns a citation bonus.
I would treat schema as a way to describe content, not as a substitute for writing it. If the most useful product detail sits only inside a script, the visible page is still failing the reader who came looking for that detail. If it appears in the body copy, the reader can see it and the bot has a passage to extract. Neither choice promises a citation. Only one gives the page a usable answer.
Structured data still has a job in traditional search. But if your answer exists only in a script, you have made it harder for a human reader to use and have not established a citation advantage. Put the fact in clear, visible body copy first. Do not confuse a machine-readable label with an answer worth quoting.
Myth 3: “An llms.txt file gets you into answers”
The appeal of /llms.txt is obvious: one Markdown file at the root of your domain could tell assistants what your company does and which pages matter. It sounds much easier than making those pages useful.
The fetch data is less encouraging. In an Ahrefs study of 137,000 domains with an /llms.txt file, 97% of those files were never fetched by AI answer engines. Of the traffic that did reach them, roughly 12% came from third-party audit crawlers, vulnerability scanners, and prompt-injection researchers. SE Ranking’s analysis of 300,000 domains found no correlation between having the file and being cited by AI; removing it from their predictive model improved accuracy.
The distinction between publishing a file and having it used in an answer is the whole argument. You can make an impeccable index of pages. If the answer engine never requests that index, it changes nothing about the live fetch. Even if it does request the file, the relevant page still needs to contain something more useful than a link to another page. A directory is not the destination.
A live answer needs material relevant to the user’s question. A site-wide index is not a replacement for that material, and publishing one does not mean a retrieval bot will read it. Put the answer on the page the bot might fetch. If you are unsure whether it can read that page at all, start with a 45-minute fetchability check, not a new file at the domain root.
Myth 4: “Category pages carry your authority into AI answers”
For conventional SEO, a category hub can be valuable. You gather related services, link to deeper pages, and give visitors a way to browse. It is tempting to assume the same hub will become an assistant’s definitive source on the subject.
Our logs gave us a blunt counterexample. Across the same 30-day window, our category and landing pages received zero live-answer retrieval hits. Our single-topic guide to YouTube ad frequency and policy changes received 273 from OpenAI and Perplexity user bots. Those counts show where the bots went, not whether an assistant cited the guide in its final answer. But a hub that drew no live-answer fetches in that window had no chance to supply a passage through those fetches.

A hub often works by sending a person somewhere else: service cards, short teasers, logos, and Learn More buttons. A retrieval bot looking for a usable passage may find little explanation in that layout. The focused guide has a better chance because the relevant details live together on one URL. That does not make category pages useless. It means they have a different job.
This is where the old SEO instinct can send a team in circles. The hub looks important in a site map, so the team polishes its headline, adds another card, and waits for it to become the source. None of that puts the explanation on the hub. If the page’s purpose is to route visitors to a detailed guide, let it do that. Judge the guide by whether it answers the question, rather than asking the navigation page to do two jobs badly.
Keep the hub for navigation; give the answer its own page. If a reader has to follow another link to find the substance, the hub is a weak candidate for a citation about that substance.
Myth 5: “Blocking AI bots is one decision”
When a company worries about training crawlers, Disallow: / can look like a clean answer. When it wants AI visibility, leaving everything open can look equally clean. Neither approach distinguishes what the bots are doing.

OpenAI documents separate roles: GPTBot for training, OAI-SearchBot for search indexing, and ChatGPT-User for on-demand fetches. Anthropic likewise distinguishes ClaudeBot, Claude-SearchBot, and Claude-User. Blocking a training crawler is not the same choice as blocking a bot that fetches a page during a user’s query.
Perplexity adds another wrinkle. Its documentation says PerplexityBot indexes pages and honors robots directives, while Perplexity-User handles user-initiated fetches and generally ignores robots.txt rules. That is one reason a single file cannot stand in for a considered access policy. And when Cloudflare measured AI crawler traffic, it found very different crawl-to-referral ratios: near 50,000 to 1 for Anthropic and 118 to 1 for Perplexity. Those are not interchangeable kinds of traffic.
I would separate two questions before touching a rule: which bot is making the request, and what response does it receive? The first tells you whether you are looking at training, indexing, or an on-demand fetch. The second tells you whether your intended policy is what the bot actually encounters. A tidy robots.txt file is not much comfort if an edge challenge answers the request instead of your page. Equally, one user-initiated fetch is not proof that you chose to permit every other kind of crawl.
Decide which activity you want to permit, then check the rules and responses for each bot. If you need a starting point for edge configuration, our AI crawler swipe file lays out copy-ready blocks. A blanket rule is easy to write. Living with the wrong one is less fun.
Myth 6: “Perplexity and ChatGPT reward the same pages”
One AI visibility bucket makes for a simpler reporting slide. It also hides differences in the pages the bots fetch. In our 30-day logs, Perplexity-User repeatedly fetched structured tables and dated change records, while ChatGPT-User concentrated on broader conceptual explainers. Those are observations from our site, not a universal scoring formula. They also describe fetches, not a promise that either format earns a citation.
The useful question is narrower than “What content does AI like?” Look at the URLs each user bot requests and what those pages contain. If the patterns differ, a single dashboard total can conceal the difference you need to act on. Inspect which bot fetches which URL before assuming one successful page format works everywhere.
Myth 7: “Monitoring your mentions improves them”
This is the myth I find hardest to kill. A tool tracks your brand across fifty synthetic prompts, a dashboard plots citation share, and everyone in the meeting gets to feel that something happened. I have no quarrel with measuring an outcome. I have a quarrel with mistaking the measurement for the work.
A dashboard is a thermometer, not a heater. If it says your share on a query fell from 14% to 11%, it has not changed a server response, fixed a blocked edge rule, or written a passage an assistant can extract. In paid search, I would not celebrate a report showing conversions fell while leaving the bids and landing pages untouched. Calling the same habit AI visibility strategy does not improve it.
The report can still tell you where to look. It cannot tell you, on its own, whether the bot was blocked, fetched the wrong page, or reached a page with no usable answer. Those are different problems. Treating them as one falling line produces the usual meeting outcome: another meeting about the line. I would rather open the logs and the page.
If you want to act on what you measure, do the work in this order:
- Check fetchability at the edge. Inspect your Cloudflare or server logs for
Claude-User,ChatGPT-User, andPerplexity-User. Look for HTTP 200 responses rather than 403 blocks or challenge screens. A page an on-demand bot cannot fetch is not much of a citation candidate. Check the actual response before deciding that the content needs a rewrite. - Build dedicated answer pages. Pick a question your prospects ask. Skip the 500-word introduction and put a self-contained factual answer under the primary heading. Give the bot something useful on that URL, not a promotional brochure pointing elsewhere. Then read the answer without the rest of the site around it: does it still make sense?
- Verify with server logs. Check weekly which URLs live user-agents fetch. A synthetic prompt run can show you an answer; your logs show whether a bot reached your pages. Keep those observations separate. A fetch is evidence of access, not proof of a citation.
Getting cited is not an outcome you can monitor into existence. At groas, our operating philosophy is that recommendations are cheap and execution is the product. Build pages bots can parse, fix the access problems, and keep checking the work. Or keep paying to admire the graph of someone else getting cited.

