ChatGPT recommends three of your competitors and leaves you out. Before anyone commissions 20 blog posts, check whether the answers you already published can get through the door.

A missing citation in ChatGPT, Perplexity, Gemini, or Google AI Overviews can point to four different failures: no page answers the question, a page exists but bots cannot fetch it, bots can fetch it but cannot pull a useful answer from it, or a third-party source wins the citation instead. Only the first is a reason to create a new page. The rest call for technical work, a better presentation of existing facts, or off-site authority. Diagnose the failure before you buy more words.

Why shouldn’t I commission 20 articles first?

Because an editorial calendar cannot fix a locked doorway. If a firewall drops a retrieval bot at the edge, or a comparison table appears only after client-side JavaScript runs, another 1,500-word glossary does nothing. You could spend months answering questions your site already answers, on pages the engine still cannot read.

I understand why content is the reflex. Marketing can assign a writer faster than it can get an engineer, an editor, and whoever controls the CDN into the same conversation. Content shops also sell content. Neither is a good reason to mistake every missing citation for an unwritten article. Audit the path from question to citation before assigning the work.

Step 1: Which questions do buyers actually ask?

Skip the vanity prompt: “What does [Brand X] do?” A buyer who asks that already knows you exist. Your useful test set comes from comparative, problem-oriented questions such as “What is the best enterprise fleet tracking software for cold-chain logistics?” or “How much does a commercial HVAC retrofitting cost per square foot in Ohio?”

Build a list of 30 to 50 prompts from sales calls, lost-deal CRM notes, and organic search term reports. Keep the wording close to what buyers used rather than polishing every question into marketing language. If the list gets unwieldy, use the prompt-list rant to narrow it to the questions that could actually cost you revenue.

Start with buyer questions, not questions about your brand. Otherwise, you are measuring whether the engine recognizes your name, not whether it recommends you when a customer needs an answer.

Step 2: Who gets cited instead of me?

Run each prompt across ChatGPT Search, Perplexity, Google AI Overviews, and Claude. A single response is not a baseline: use fresh sessions or an API to limit personalization, following the baseline measurement. For each answer, record:

  • Whether your brand appears in the response.
  • The exact URLs in its footnotes or inline source chips.

A clipboard audit sheet comparing AI search engines, with checklist rows and diagnostic stamps.

Then look at the source that got the slot. Was it a competitor’s feature page, an independent review, a trade journal, or a two-year-old Reddit thread? That distinction matters. A competitor page suggests one kind of contest; a discussion or directory suggests another. Record the winning URL alongside the lost prompt rather than writing “not visible” in a spreadsheet and calling the audit done.

The log gives you a practical next question: Did your page fail to exist, fail to load, fail to answer plainly, or lose to another source? That is what you classify next.

Step 3: Which of the four gaps caused the miss?

Take each uncited buyer question through these checks in order. Do not assign a writer until you have looked for an existing answer and checked whether a retrieval bot can reach it.

Gap 1: Do I have a page that answers it?

This is the literal content gap. A buyer asks how your cold-chain tracking hardware handles sub-zero calibration drifts, and your site has never published the answer: not in a spec sheet, a knowledge base article, or a pricing matrix. You cannot expect an engine to cite an answer your domain does not contain.

In that case, write the page. Make it answer the question rather than merely mention the topic. But first search what you already own. Teams often have the relevant detail buried in a service page or technical document and assume they need a new URL because that page never earned a citation. No existing answer means write; an existing answer means keep diagnosing.

Gap 2: Can a bot fetch the page that exists?

A page can look complete in your browser and still deliver little to a retrieval bot. Two places to check are client-side rendering and edge security.

If a comparison table, pricing calculator, or feature matrix appears only after React or Vue runs in the browser, inspect the raw HTML returned before that code executes. If it contains an empty container where the answer should be, a crawler fetching that HTML has little to use. The discussion of JavaScript and AI crawlers is a useful starting point, but test the page you actually serve rather than assuming the browser view tells the whole story.

Your web application firewall may be the other obstruction. Research across 1,444 domains found server-level AI crawler blocks; firewall rules can return HTTP 403, 520, or 521 responses. Check your own event and access logs for the affected requests. A user-agent string alone does not establish that a requester is legitimate: scrapers spoof them, which is why bot verification methods matter when changing rules. This account of ChatGPT-User blocks illustrates the kind of edge failure worth investigating.

If the answer is trapped behind rendering or a block, fix delivery before rewriting it.

Gap 3: Is there a sentence worth quoting?

Suppose the page loads cleanly and the answer is present, but it reads like this: “Our state-of-the-art platform delivers comprehensive assurance for perishable logistics.” That may satisfy a brand review. It does not tell a reader which standard applies, what the system measures, or how it handles a failure.

The Princeton GEO study by Aggarwal and colleagues found visibility gains from changes including direct quotes and statistical evidence. The operational lesson is simpler than the research terminology: put verifiable, specific answers where a retrieval system can extract them. For an enterprise cold-chain compliance question, that could mean the temperature thresholds, regulatory standard codes, documented failure percentages, or named protocols already supported by your materials. Do not invent a statistic to make a paragraph look quotable.

This is editing, not necessarily content production. Replace the vague benefit blurb with the concrete answer your business can substantiate; use a table where the comparison genuinely needs one. A fetched page still has to say something an engine can use.

Gap 4: Does a third-party source win anyway?

This is the maddening one. Your page loads, gives a clear answer, and still loses the citation to a Reddit thread or comparison directory. For questions asking what buyers think or which option they should trust, a first-party product page may not be the source the engine chooses. The reported prominence of Reddit citations in generative search is a reminder that the contest extends well beyond your own domain.

A buyer asking “Which CRM is actually best for a 50-person agency?” may get an answer assembled from peer discussion, reviews, and trade coverage rather than vendor copy. Traditional rank is no guarantee either: data on Google AI Overview citations challenges the assumption that a top-ten organic result must win an AI citation.

Do not keep polishing a feature page as though one more paragraph will replace independent discussion. Check which outside sources won, then decide where your business needs an accurate presence. If the winning answer is off-site, the work may need to be off-site too.

Diagram of the four-gap AI citation audit: missing pages, fetchability, quotability, and third-party authority.

Step 4: What should I fix first?

For execution, I would usually work in this order: 2 → 3 → 4 → 1. Classification and prioritization are different jobs. You check whether a page exists to identify the gap; you fix shared delivery problems early because one blocked template or rule can affect many existing answers.

  1. Fix fetchability (Gap 2). Inspect raw HTML and the relevant firewall events. Correct the rendering or verified-bot access problem you find. There is little point improving copy a retrieval bot cannot reach.
  2. Make existing pages quotable (Gap 3). Replace empty claims with the specific answers, tables, and documented parameters you already have. Work on the URLs that address your high-intent prompts before opening new drafts.
  3. Build an accurate off-site presence (Gap 4). Where outside sources repeatedly win, participate candidly in relevant discussions and keep directory or review profiles accurate. Do not treat a vendor-page rewrite as a substitute for third-party context.
  4. Write what is genuinely missing (Gap 1). Commission a new page when your audit confirms that no existing one answers an important buyer question.

This is a work order, not an excuse to leave a known missing answer unwritten while a developer clears an unrelated ticket. Prioritize the gaps tied to valuable prompts. Skip net-new content when an existing page is blocked or merely vague.

Editorial illustration of a locked server-room door beside a stack of ignored marketing brochures.

Step 5: What should an AI visibility platform actually show me?

A prompt-monitoring dashboard can tell you where you lost. That is useful, but it is not the same as telling you why. If you are evaluating platforms such as Profound, Peec AI, or Otterly.ai, look past Share of Voice and sentiment charts. A platform comparison is a place to start; the decisive test is whether a tool can connect a lost prompt to the failure mode you need to fix.

Before paying $500 to $2,000 a month, ask the sales team:

  • Can you distinguish an HTTP 403 firewall block from a page that has not been indexed? An alert without that distinction still leaves your team doing the diagnosis.
  • Do you inspect the raw HTML available to crawlers, or only the page after a browser renders it? A rendered page can conceal an empty initial response.
  • Who makes the technical and content changes after the report arrives? A recommendation that becomes another backlog row is not execution.

That last question is why I favor groas over another monitoring-only workflow. Its autonomous growth engine is built around continuous paid and organic search execution, with a named human strategist setting direction and guardrails, rather than asking a marketer to interpret another dashboard on Friday afternoon. That does not mean you should assume any platform has access to change your firewall. It means you should make ownership explicit: who investigates the block, who can change the rule, who edits the answer, and who checks whether citations move afterward.

Buy a path from diagnosis to action, not a prettier list of misses.

Step 6: How often should I re-check, and what counts as progress?

Do not refresh your prompt panel every morning. Generative engines can vary their responses, and daily checks invite panic over noise. Use a fixed biweekly cadence: run each query ten times in a controlled batch, record citation win rates, and compare them with your original baseline.

Look for changes you can inspect, not just a proprietary visibility score. Do the logs still show 403 responses for verified retrieval crawlers? Does your domain earn citation chips on high-intent buyer prompts that previously cited competitors? Does analytics show attributable referral traffic from LLM web sessions? Those signals connect the repair to an outcome. A screenshot of one favorable answer does not.

Keep the prompt set and measurement routine steady enough to tell progress from variation.

Should I do the audit in-house or hand it off?

Do it in-house if a web developer can inspect CDN settings and server responses, and an editor can replace corporate adjectives with clear facts without waiting for an executive committee. A spreadsheet is enough to hold the prompt list, cited URLs, gap classifications, owners, and repeat results. You do not need an enterprise subscription to discover that a page is blank in raw HTML.

Hand it off if the work keeps stalling between marketing, IT, and whoever owns search. A firewall change that sits in ticket ping-pong for six weeks is not a strategy. Nor is adding manual prompt tracking to a team already juggling complex Google Ads campaigns, feed restructuring, and organic search across multiple products. Look for continuous execution with human strategic oversight, not an agency billing hours to hand back another list of article ideas.

If you take one action today, open your server access logs or firewall event log, not a new Google Doc. Check requests associated with ChatGPT-User, ClaudeBot, and Perplexity-User, and investigate any run of HTTP 403 or 521 responses. Verify the requester before changing a security rule. If legitimate retrieval requests are being blocked, fix that doorway first. Then see what the content you already wrote can do.