October 2, 2026
•
11
min read

How to Measure AI Visibility Without Tracking Fake Rankings

Young man with curly hair wearing a black shirt outdoors against green foliage background.


Alexander Perleman
, Head Of Product @ groas
Ex-Goldman Sachs and Stanford Computer Science

Email: alex@groas.com

LinkedIn: https://www.linkedin.com/in/alexander-433793253/
Cover image for: How to Measure AI Visibility Without Tracking Fake Rankings

You can find out whether AI answers put your brand in front of buyers without paying for a pretend position-4 ranking. That ranking is the mistake: run the same buying prompt three times and you can get three different shortlists, in three different orders.

 

In one test of 12 prompts run about 3,000 times, researchers found the identical list came back less than 1 time in 100, and the identical order about 1 in 1,000. I used to tell clients to watch average position for everything. I was wrong about this one.

 

There is no dependable position 1 in an AI answer. Google even counts an entire AI Overview as one block at position 1 in Search Console, so a cited card at the top and a link buried behind Show More report the same. This guide gives you an in-house measurement plan built around three things: mention rate across fixed buyer prompts, citation share, and evidence that AI crawlers fetched your pages.

 

Why can’t I track AI rank like a search result?

I treat an AI answer like a vote count, not a leaderboard. One lab test got 80 different completions in 1,000 runs even at temperature zero. In one July 2026 week, the median cited domain authority in ChatGPT jumped from 35 to 51. That is not a tidy ranking shifting one place. The set of sources showing up changed.

 

In the 3,000-run test, visibility percentage was more useful than exact placement: some brands appeared in 60 to 90% of answers even as their spots bounced around. Track the percentage, not the position. A single answer can show you what happened once. Repeated answers to the same questions tell you whether your brand shows up often enough to matter.

 

Search Console makes the distinction harder to ignore. It logs every default AI Overview link on load, while links behind toggles count only after expansion, yet they report inside the same position-1 block. Start with presence: were you cited, yes or no? Then look at what happened next: site visits, inquiries, closed business. The display position between those points is diagnostic, not proof.

 

Which three numbers should I put on the sheet?

Keep the measures separate. They answer different questions, and a blended AI visibility score makes it too easy to miss the one you can act on.

 

  1. Mention rate: how often are you named? Pick 30 buyer questions, run each three times, and count the answers that name your brand. If 36 of 90 answers mention you, your mention rate is 40%. Practitioners who run 20 to 30 industry questions across ChatGPT, Gemini, Perplexity and AI Overviews also log who else gets named and whether the tone is positive. Use the same prompts next week, or the percentage stops being a comparison.
  2. Citation share: how often is your page linked? Record a citation separately from a mention. A mention puts your name in prose; a citation gives a reader a linked page to click. For this test, calculate the share of eligible answers that link to one of your pages, using the same answer count as the mention-rate calculation. Label it clearly as share of answers citing us, rather than quietly switching to a share of all links. The gap matters: in 20.5M citations studied from March to July 2026, 90.4% named no tracked brand and only 4.5% pointed to the brand’s own page. Two columns beat one impressive-looking score.
  3. Crawler evidence: did an AI bot fetch the page? OpenAI describes three agents with different jobs: GPTBot crawls for training, OAI-SearchBot serves search discovery, and ChatGPT-User can fetch a page when a user asks about it. You can allow OAI-SearchBot while disallowing GPTBot. A fetch does not promise a citation, but it tells you whether access is part of the problem before you rewrite the page.

The order matters when you diagnose a miss. If a prompt does not mention you, check whether a relevant page is accessible before treating the copy as the culprit. If it mentions you but links elsewhere, investigate the evidence the answer chose. Do not let a mention stand in for a citation, or a crawler visit stand in for either.

 

Cartoon server log with one verified crawler and two impostors filtered out

A log line only counts if you verify it. Any script can claim to be GPTBot in its user agent, so match the IP against OpenAI’s published JSON ranges at openai.com/gptbot.json, searchbot.json and chatgpt-user.json instead of trusting the name. One July 2026 check listed 21 IPv4 CIDRs for GPTBot and no reverse-DNS path like Googlebot has. Outside-range hits are not proof of an OpenAI fetch.

 

Check your CDN rules, too. Cloudflare splits Search from Agent from Training, and from September 15, 2026, multi-purpose crawlers combining Search and Training get blocked for customers who block Training. If your aim is search access without training access, inspect the specific settings instead of assuming a blunt Block AI Bots switch does what you intend. Unverified bot charts are fan fiction.

 

How do I build a prompt set buyers would actually use?

Start from buyer questions, not a keyword export. I pull questions from sales calls, support tickets and the People Also Ask boxes customers actually click. The method is to build 30 to 100 real questions, mix category, comparison and problem prompts, tag each by intent, and fix location before running them. Keywords tell you how people search. Questions get closer to how they buy.

 

Sort the questions into three buckets:

 

  • Branded: the buyer asks about you by name. These tell you whether the answer can describe a brand it has been handed.
  • Unbranded: the buyer describes a problem or category without naming a provider. These show whether you enter a shortlist without prompting.
  • Comparison: the buyer names you and a competitor, in the is [brand] or [competitor] better shape practitioners use in a 20-to-30-question baseline. These show what evidence the answer puts beside each option.

Most small businesses overstock the branded bucket because it flatters them, then wonder why AI Overviews never mention them elsewhere. If half your set is not unbranded, rebuild it. Keep the wording fixed once testing begins. Editing a question midstream may be sensible for a future test, but it does not belong in the same trend line.

 

Photograph of sales-call questions on sticky notes sorted into branded, unbranded and comparison piles

How do I run the in-house test without fooling myself?

Question: Do buyers meet my brand inside AI answers, or only on Google page one?

 

Setup: Take 30 prompts from the three buckets above. Lock the location and logged-out state, then run each prompt three times in ChatGPT, Gemini and Google AI Overviews. That gives you 90 planned checks per model, 270 in all. Repeat the set once a week in the same order. If a check does not produce an AI answer, record that instead of inventing a mention or citation result. Keep the denominator of answers you actually measured visible beside every rate.

 

Controls: Use the same prompt wording, location, model and collection routine for each weekly set. Read results model by model before you combine anything. The point is not to make answers stop varying; it is to stop your own procedure from adding avoidable variation.

 

Record five fields for every answer:

 

  1. Mentioned: yes or no for your brand.
  2. Cited: yes or no, plus the exact URL linked.
  3. Competitors: the names that appear alongside you.
  4. Tone: positive, neutral or negative.
  5. Answer type: direct recommendation, comparison table, or how-to with sources at the bottom.

Expectation: I expect unbranded mention rates to swing more than branded rates, potentially 15 to 20 points week to week. Unbranded answers have a wider pool of possible recommendations, and each run can draw a different shortlist. Treat that as an expectation to test, not a correction to apply to your sheet.

 

What to measure: Calculate mention rate and the share of answers citing your page separately for each model and prompt bucket. Keep the raw answers and URLs. A weekly total is useful for scanning, but a prompt where a competitor repeatedly gets cited tells you more about what to fix than a blended score does. One practitioner method runs the same set per engine weekly and reports rates rather than single checks. If ChatGPT cites you in 50% of its answers while AI Overviews cites you in 10%, their average hides the work.

 

Your week-one result is a baseline with visible variation, not a verdict. Resist the urge to rewrite a page because one answer left you out.

 

When is a change large enough to act on?

One bad week means little. In practitioner baselines, only about 30% of brands stay visible consecutively and 1 in 5 holds visibility across five runs. My working rule is three weekly sets for a baseline, then action when a prompt group moves the same way twice in a row.

 

A competitor beating you in one of three runs is a reason to watch. A competitor cited in seven of nine runs while you appear in one is a gap worth inspecting. I watch a 10-point swing in one model when the others are flat; I investigate the same 10-point drop across two models. Those are decision rules for this sheet, not laws of AI search. Never change content on a single Wednesday reading.

 

What does a crawler visit tell me about a missing citation?

Separate access from selection. GPTBot activity alone does not show that ChatGPT search can use a page; check OAI-SearchBot separately. OAI-SearchBot fetches with no citations suggest that access is available but your page was not chosen in the answers you measured. Neither log pattern, by itself, explains why.

 

Before rewriting anything, run the access check in Before You Write for ChatGPT, Check Whether AI Can Read Your Site. A forgotten disallow or CDN setting is a dull explanation, but dull explanations are cheaper to fix than an unnecessary content project. Verify the fetch, then inspect the answer.

 

How do I turn the results into one useful content fix?

Sort first for prompts where competitors are cited and you are not, especially comparison prompts. Those are not automatically writing problems. The answer wanted evidence and selected another page; your job is to inspect what that page answered and what yours did not make clear.

 

That is where the fix in How to Get Cited in AI Overviews and ChatGPT: A Passage-First Guide earns its keep: a self-contained passage answering the prompt’s question with a number, a step or a definition. Do not add a paragraph merely because a dashboard turned red. Give the answer something useful to cite.

 

Second, sort for pages AI fetches but your sampled answers never cite. I see this with pricing and versus pages that ramble for 2,000 words without one quotable block. If the bot visits for four weeks and your fixed prompts still produce no citation, you have moved past the access check. Fix cited-by-competitor prompts first; fetched-never-cited pages second. Then rerun the same prompts to see whether the change helped.

 

When is a monitoring tool worth paying for?

Buy a tool when maintaining the spreadsheet costs more than the subscription. At, say, $20k a month in spend, two hours a week copying answers into sheets may be manageable until a second location or a fifth competitor set arrives. Sampling at scale is a real job. A decorative rank is not.

 

Ask a vendor to show you the prompt set, the answer samples, the linked URLs and the denominator behind each rate. Then ask whether it can connect verified fetch data to the pages you are investigating and show a before-and-after on the same fixed prompts. Monitoring alone cannot tell you whether OAI-SearchBot was allowed to fetch a page or what changed after you rewrote it. If the product sells a rank without those checks, skip it. I would pay for repeatable sampling and verified fetch data tied to the same questions.

 

What should I do in the first month?

  1. Week 1: establish the baseline. Lock 30 prompts, run the three checks per prompt and model, and record mention and citation rates separately. Change nothing.
  2. Week 2: check access. Verify crawler requests, inspect blocks and confirm whether OAI-SearchBot can fetch the pages you care about. Change no copy yet.
  3. Week 3: make two targeted edits. Rewrite one fetched-never-cited page and one comparison passage competitors keep winning. Give each a number or step the answer can use.
  4. Week 4: rerun the fixed set. Compare rates model by model and prompt by prompt. Keep watching for a repeatable move rather than declaring victory from one answer.

Start today by writing down 10 unbranded questions your last 10 buyers asked before they bought. If you cannot list 10 from memory, pull them from call notes before you open any tool.

Veelgestelde vragen

Kun je de ranking van je merk in AI-antwoorden volgen zoals bij Google?

Nee. Een zelfde koopvraag kan drie keer achter elkaar verschillende shortlists in andere volgorde opleveren, en hetzelfde merk kan verschijnen in 60 tot 90 procent van de antwoorden terwijl zijn plek steeds wisselt. Meet daarom hoe vaak je merk verschijnt als percentage, niet op welke plaats het staat.

Welke meetwaarden gebruik ik om zichtbaarheid in AI-antwoorden te meten?

Houd er drie apart: een mentieratio, oftewel hoeveel antwoorden je merkvraag benoemen; een citatieshare, oftewel welk deel van die antwoorden naar jouw pagina linkt; en bewijs dat een AI-bot daadwerkelijk toegang had tot je pagina's. Meng ze niet tot één score, want dan mis je waar je kunt ingrijpen

Wat voor vragen moet ik gebruiken als vaste promptset?

Begin bij echte kopersvragen verkregen via verkoopgesprekken, supportscores en People Also Ask-boxes. Mix gebrandeerde, ongebrandeerde en vergelijkingsprompts, label elk op bedoeling en houd locatie constant. Zorg dat ten minste de helft ongebrandeed is, anders overgeef je misschien vooral wat goed voelt maar geen kortetermijnuitkomsten toont. Verander pas tekst na afloop van een testsessie.

Hoe weet ik zeker dat een logregel echt afkomstig is van GPTBot?

Match het bronbestand blokkerendje Domeinnaam ranges