Say a buyer asks ChatGPT for the most reliable emergency plumber in Austin for trenchless sewer repair. Your service page has a license number, trenchless equipment, a 4.8-star average across 600 reviews, and same-day dispatch. ChatGPT cites a competitor with a thinner page. It may never have compared the two.
On our own edge logs, live-answer bots hit robots.txt more than any real page. That is a strange place to start a discussion about citations, but it tells you where to look: getting cited is not simply winning a ranking. It is surviving a retrieval pipeline. A buyer’s question becomes several searches; a bot fetches candidates; access rules and page rendering determine what it can read; then the assistant works with passages, not your lovingly arranged service page. I spent nine years inside Google Ads accounts, and I keep seeing owners treat a missing AI citation as a content-quality verdict. Often it is a mechanical dropout. Let’s follow that plumber question through the gates in order.
First principle: a citation is the end of a sequence
In Google, a plumber page can sit at position four and still get clicks. In a ChatGPT answer that names two or three shops, an unmentioned shop gets no link from that answer. A good page can disappear before the assistant has enough text to judge it.
Here is the working model, in order:
- Query rewriting: the buyer’s question becomes smaller, retrieval-ready searches.
- Live fetching: a bot requests candidate pages.
- Access: robots rules, site security, and the returned HTML determine what the bot can read.
- Passage extraction: usable text becomes short candidate answers.
- Source selection: the assistant chooses passages it can use and names sources.
The order matters. Improving a cost paragraph cannot help a fetcher that never reaches the page. Fix the earliest gate where you drop out, then move forward.
Stage 1: your buyer’s question becomes several searches
The plumber question is not necessarily searched as one long sentence. ChatGPT can strip filler, pull out entities and constraints, and split a hard question into sub-questions. One extended test found single buyer questions fanning out into 5 to 10+ rounds of searches per answer, including site: probes aimed at domains the model already trusts.
For our plumber, plausible searches include:
- trenchless sewer repair Austin
- emergency plumber Austin same day
- Texas plumber license lookup
- trenchless vs excavation cost
- best reviewed trenchless plumbers Austin
The failure mode is easy to miss. Your page says, “We are the most reliable emergency plumber in Austin.” The retriever may be looking for a license, a method, a dispatch window, or a cost explanation. A broad promise is not an answer to those smaller questions.
Put a short, plain answer where each relevant question belongs: methods on the trenchless page, the license number in readable text, the dispatch window on the emergency-service page, and service areas stated by neighborhood. I used to tell clients one strong service page was enough. I was wrong. One strong page without retrievable sub-answers can look irrelevant to a very specific search.
Stage 2: a live bot fetches candidate pages
Two kinds of AI bots do different jobs. Index crawlers help build the pool of pages an engine can surface. Live fetchers request content around answer time because a person has asked a question. In a honeypot test, ChatGPT-User fetched selected page content after the search step, rather than OAI-SearchBot. Perplexity draws a similar distinction: PerplexityBot builds its index and honors robots.txt, while Perplexity-User fetches in response to a person’s request.
That distinction changes what a block does. Block an index crawler and you may affect future surfacing. Block a live fetcher and you may prevent a page from being read for an answer. Calling both of them “the AI bot” hides the decision you actually made.
It also explains that odd line in our logs: live-answer bots requested robots.txt more than any real page. They were checking the rules before requesting content. A plumber who blocks bots to “protect content” needs to know which bots the rule covers. The AI crawler swipe file lays out the split between training crawlers and live-answer fetchers.
The practical question is not whether your site is crawlable in some general sense. It is whether the fetcher that arrives for a buyer’s question can reach the relevant page.
Stage 3: access determines what the fetcher can read
Reaching the URL is not the same as reading the service page you see in a browser. The fetcher can encounter robots rules, a CDN or security plugin that filters bots, and HTML that contains little of the visible page. A service page may look complete to a customer while its important details arrive later through JavaScript.

That creates an awkward split. Your reviews load in a widget, your price list sits behind tabs, and your city pages render client-side. Google may render content that a live-answer fetcher does not. These fetchers can read the initial HTML and move on without executing JavaScript. A page that performs well in conventional search can therefore give an answer-time fetcher very little to work with.
Robots rules deserve the same precision. Blocking PerplexityBot can remove future Perplexity surfacing without serving as a training opt-out; user-triggered fetches can still read public pages. An index block and a readable live page are different things. So are a successful HTTP request and useful page text.
I check the raw response with curl rather than trusting a dashboard preview, then compare it with what the browser displays. Curl alone will not tell you how every bot was treated by your site security, but it will expose a common failure fast. If your license number, service area, and dispatch promise are missing from the initial HTML as plain text, fix that before polishing copy for citations.
Stage 4: the assistant needs a passage, not an entire guide
Now suppose the plumber page clears the fetch and access checks. The assistant still has to find text that answers one of its smaller questions. A 3,000-word guide to trenchless sewer repair may contain everything a customer needs, yet make a particular cost or dispatch answer hard to extract. A short block under a precise heading gives the assistant a cleaner candidate.
That is not a rule against long pages. It is an argument for self-contained passages inside them. The passage should make sense without the paragraph before it. If a buyer copied it into a message, would the recipient know what service, location, and claim it describes?
In a Princeton and Georgia Tech test, adding statistics lifted AI visibility about 41%, while citations and quotations lifted it 30 to 40%. That does not mean sprinkling numbers onto a plumber page. Numbers help when they answer the question and belong to the business making the claim. Invented precision is still invented.
For our illustrative plumber, assume the prices, timing, and license below are accurate for that shop. A passage might read: “Trenchless sewer repair in Austin costs $135 to $185 per foot for pipe bursting. Most jobs finish in one day, and we dispatch same-day across central Austin under Texas license M-41208.” It gives a reader one bounded answer rather than a slogan. The figures and license are part of the hypothetical example, not a market quote or a real shop’s credentials.
Put that kind of block under a heading that names the question, in readable HTML. Repeat the pattern for method, license, and dispatch where those details are true. If your competitor offers a quotable answer and you offer only a long introduction, the thinner page has an advantage at this gate.

Stage 5: the assistant selects a source it can support
Now two plumbers have cleared the mechanical gates. Both pages load; both contain a useful cost passage. The assistant still has to choose what to cite. This is where corroboration becomes useful: does each shop look like the same business wherever else it appears? Check the legal name, license number, service area, and phone format on the site, Google profile, Yelp page, and relevant local directories.
Consistency is not a guaranteed tiebreaker. It does make a passage easier to connect to a recognizable business than five conflicting descriptions do. Being mentioned and being cited are different outcomes. A shop can be known to an engine and still lack the passage it chooses to link.
One test of 1,000 B2B questions found only 0.09% of pages cited by all three linking engines, while pairs of engines named the same vendors 35 to 42% of the time. The useful distinction for our plumber is between recognizing a business and selecting a particular page as proof. Each engine can reach a similar business choice through different sources.
Conventional ranking still matters, but it does not settle the choice. In that test, 74% of cited pages sat outside Google’s top 10, with 12 to 38% overlap. Position one was quoted at 86.9% in Perplexity, against 46% at position nine. Engine habits differed, too: Perplexity cited Reddit 1,051 times on those questions, while ChatGPT and Claude cited it zero times.
Do not read those figures as a new ranking formula. Make the business identifiable, then give the engine a readable passage that supports the answer. A familiar name without proof and a perfect passage attached to muddled business details can each lose for a different reason.
Find your first failing gate before you rewrite anything
Run these checks in order. Stop at the first failure, fix it, then run them again. Otherwise, you risk improving a passage no fetcher can read.
- Rewrite check: Take the buyer’s question and list the likely sub-questions, as we did for the plumber. You can ask ChatGPT to suggest searches, but do not mistake its suggestions for a record of searches it actually ran. Does your page answer the relevant questions in plain text?
- Fetch check: Look in server logs for visits from ChatGPT-User, Claude-User, and Perplexity-User to the money page. No visit in the last 30 days does not prove you were never considered. It does mean you should test whether the page can be fetched before blaming its prose.
- Access check: Request the page with curl and read the returned HTML. If the price, license number, or service area appears only after JavaScript runs, the fetcher may not see the details you are relying on.
- Passage check: Find a short block a buyer could quote without editing. Does it answer a specific question on its own, or does it need three paragraphs of setup?
- Entity check: Search the exact business name alongside the license number. Where your site, profiles, and directories disagree, correct the underlying business details.
The first failure is the place to work. That is more useful than a general instruction to “optimize for AI,” which I usually translate as “do several jobs without knowing which one is broken.”
Try the model on a different buyer question: “What is the best CRM for a 20-person law firm with conflict checks?” The likely sub-questions concern conflict-check features, legal CRM pricing per seat, migration effort, and bar compliance notes. Your CRM page says “trusted by law firms” and shows three logos. It does not answer those questions, so start at Stage 1. Do not commission a better testimonial yet.
Suppose you add clear answers and the page becomes a plausible candidate. Pricing still lives inside an interactive calculator: check the initial HTML at Stage 3. Suppose that is readable, but the only explanation of conflict checks runs for 400 words without a self-contained answer: work on Stage 4. Finally, compare the product’s name, category, and review count across G2, Capterra, and its own site for Stage 5. The same sequence diagnoses a plumber or a CRM without pretending they need the same content.
For teams that clear the mechanical gates, I keep the passage-first playbook for earning citations close. Not every site should start citation work: thin fulfillment, no reviews, or inconsistent business details are more basic problems. For everyone else, clear the gates in order. Let the competitor keep polishing the guide nobody can fetch or quote.
Frequently asked questions
Why does ChatGPT cite a competitor with a worse page than mine?
A missing citation is often a mechanical dropout, not a content-quality verdict. The buyer's question becomes several searches, a bot fetches candidate pages, access rules and HTML determine what it can read, then the assistant works with short passages and chooses sources. A good page can disappear at any gate before the assistant has enough text to judge it.
How does ChatGPT turn a buyer's question into searches?
ChatGPT rewrites the question into smaller, retrieval-ready searches, often 5 to 10 or more rounds per answer, including site: probes on domains it already trusts. A broad promise like being the most reliable plumber is not an answer to those smaller questions about a license, method, dispatch window, or cost. Put short plain answers where each relevant question belongs.
What is the difference between ChatGPT-User and OAI-SearchBot?
Index crawlers like OAI-SearchBot help build the pool of pages an engine can surface, while live fetchers like ChatGPT-User request content at answer time when a person has asked a question. Blocking an index crawler may affect future surfacing, but blocking a live fetcher can prevent a page from being read for an answer. Perplexity draws a similar line between PerplexityBot and Perplexity-User.
Do AI answer bots see content loaded by JavaScript?
Often not. Live-answer fetchers can read the initial HTML and move on without executing JavaScript, so reviews in a widget, prices behind tabs, or city pages rendered client-side may be invisible to them. Check the raw response with curl and compare it with what the browser displays; if key details are missing from the initial HTML, fix that before polishing copy.
What makes a passage easy for ChatGPT to cite?
A short, self-contained block under a heading that names the question, which makes sense without the paragraph before it. It should give one bounded answer naming the service, location, and claim, such as a cost per foot, a completion time, and a license number. Adding statistics lifted AI visibility about 41% in one test, but only when the numbers answer the question and belong to the business.
Does consistent business information across the web help AI citations?
Yes, as corroboration. Consistency of the legal name, license number, service area, and phone format across the site, Google profile, Yelp, and directories is not a guaranteed tiebreaker, but it makes a passage easier to connect to a recognizable business than five conflicting descriptions do. Being mentioned and being cited remain different outcomes.
Do I need to rank first on Google to get cited by ChatGPT?
No. Conventional ranking still matters, but in one test of 1,000 B2B questions, 74% of cited pages sat outside Google's top 10, with only 0.09% of pages cited by all three linking engines. Position one was quoted 86.9% in Perplexity against 46% at position nine, so a high position helps but does not settle the choice.
Where should I start if my business is not cited by AI assistants?
Run checks in order and fix the first failure: list likely sub-questions and check your page answers them in plain text, look in server logs for visits from live fetchers like ChatGPT-User, read the raw HTML with curl, find a short quotable passage, and verify business details match across profiles and directories. Stop at the first failure before improving anything else.




