October 5, 2026
•
11
min read

Your AI Visibility Report Is a Screenshot Until It Shows What Changed

Young man with curly hair wearing a black shirt outdoors against green foliage background.


Alexander Perleman
, Head Of Product @ groas
Ex-Goldman Sachs and Stanford Computer Science

Email: alex@groas.com

LinkedIn: https://www.linkedin.com/in/alexander-433793253/
Cover image for: Your AI Visibility Report Is a Screenshot Until It Shows What Changed

I spent a decade building PPC reports at 1am that looked impressive and proved nothing. If your AI visibility report cannot tie a citation on a buyer prompt to a specific change you made, it is a screenshot collection.

 

How can my business measure brand visibility across ChatGPT, Gemini and Google AI Overviews?

Start by counting four different things. I learned this the PPC way: impression share told me I showed up; it never told me I won a click on a buying query. AI answers have the same split. Track mentions, citations, recommendations, and source appearances, but do not pretend they mean the same thing:

 

  • Mention: Your brand appears in the answer text.
  • Citation: Your URL appears among the answer’s sources.
  • Recommendation: The answer explicitly suggests your business.
  • Source appearance: Your brand appears on a third-party page the answer cites.

Mentions and recommendations are outcomes. Citations and source appearances are levers you can investigate and fix. If a cited third-party page names your competitor but not you, that tells you more than a blended visibility score ever will. Most reports roll these signals together. That is how we ended up with PPC dashboards full of green arrows and flat revenue.

 

You also need more than one answer to judge any of them. SparkToro had volunteers run 12 prompts 2,961 times and found identical brand lists in the same order about 1 in 1,000 times. One run per prompt is noise. AirOps found a similar pattern: only about 30% of brands appear in consecutive answers, and about 1 in 5 appear across five runs.

 

So record the percentage of repeated runs in which you appear, per buyer prompt and per engine. Keep the citation count beside the mention count. If ChatGPT names you but cites somebody else, that is not the same result as a buyer seeing your page in the sources. And if you appeared once on Tuesday, you have an observation, not a trend. Start with that distinction before you buy a dashboard to blur it.

 

Cartoon marketer framing one ChatGPT screenshot while buyer prompts pile up behind him

How can a small business automate brand monitoring in AI search in-house?

You can, and I think you should start there. I used to tell clients the monthly report was the hard part. I was wrong. The hard part was acting on it.

 

For a baseline, take 10 to 20 prompts, run them in clean sessions with no history across ChatGPT, Perplexity, Gemini and AI Overviews, then score mentions, citations, recommendations and which competitors appeared. Put the answers in a spreadsheet. Run the money prompts twice. That is about an hour a month, costs nothing, and gives you most of the signal you need before you pay for a tracker.

 

The sheet needs a row for each prompt and engine, plus the date and raw answer. Otherwise you will remember that your brand appeared and forget whether it appeared in the answer or only on a cited page. You do not need a beautiful template. You need to be able to compare this month with last month without changing what you counted.

 

Where manual tracking breaks is not quite where vendors say it breaks. They will tell you it does not scale. True, but hardly the pressing problem at 20 prompts. The more immediate ways to wreck your baseline are familiar:

 

  • Nobody records which model version, location or conditions they checked.
  • Someone adds branded questions because they make the trend line look nicer.
  • The sheet identifies a missing citation, but nobody owns the page or source that needs work.

Single-platform checks can hide 60 to 70% of your footprint, while manual spot checks miss variation and offer no alert when a citation drops. Those are reasons to make your routine consistent, not reasons to buy software before you know what you are looking for.

 

Automate the reminder before you automate the running. Same prompts, same sheet, one owner, one monthly date on the calendar. If your team cannot keep that appointment, a paid tool will just help you ignore a fancier dashboard.

 

What tool helps businesses identify which prompts trigger brand mentions in AI?

No tool can rescue a prompt list built around the wrong questions. A tracker reruns the list you give it; it does not decide which buyer conversations deserve your attention. This is where most teams lose before they start.

 

AI answers respond to buyer language, not the headings in your keyword planner. If you track 30 variations of “best [category] software,” you may measure the same generic answer 30 times while missing the comparison or problem question asked just before somebody picks a vendor. Build a list of 15 to 50 buyer questions across category, comparison, alternatives, problem and branded prompts. Weight it toward commercial intent. Keep branded questions separate so they cannot inflate your unbranded trend.

 

One B2B SaaS team tracking 40 prompts in those buckets rose from an 8% to a 27% mention rate in 90 days after adding comparison pages. The useful lesson is not to copy its prompt count. It is to make the questions resemble the ones buyers ask, then work on the pages that might answer them.

 

Freeze the list once you have a baseline. I learned this PPC lesson late: the minute you add branded prompts to juice the line, the trend is dead. Keep the same prompts and engines in the same order month to month. For each prompt and engine, record mention rate, citation rate, share of voice, answer position and sentiment in a separate row rather than one blended score. You can review the raw answers when a row changes; you do not need to read every answer to discover that the average barely moved.

 

Start with the prompts that carry buying intent. If a question would not help a buyer choose, it should not take a place in your small monthly panel just because it makes the count look thorough.

 

Overhead view of a spreadsheet mapping buyer questions to AI brand mentions

What is the best software to automate AI search visibility reporting without hiring an agency?

I am going to push back on the premise, because this is where I burned clients for years. Software automates the running, not the fixing. A tracker will rerun prompts and chart a line. It will not rewrite your comparison page, earn a third-party citation or fix a page the model cannot read. Buy it expecting movement, and you have bought a faster version of my 1am PDF.

 

For tracking alone, entry pricing is public and boring in a good way. Otterly starts at $29 a month for 15 prompts with daily tracking, covering ChatGPT, AI Overviews, Perplexity and Copilot, with Gemini and AI Mode as add-ons. Peec starts around $95 a month for 50 prompts. Profound lists about $99 a month for a ChatGPT-only starter. Broader suites cost more once you add the engines you need.

 

Buy the smallest tier that covers the three engines your buyers use. Before paying, check whether it reruns prompts enough to show variation, separates engines, and saves raw answers alongside model version and location. A single blended score may look clean in a meeting. It will not tell you which buyer prompt lost its citation. If that is all a tool gives you, pass.

 

If you like plumbing, you can also build a modest checker. A Perplexity-based check runs about $0.01 per query, so a 20-prompt monthly audit comes in under $1, plus an hour to read it. Log the engine, prompt, timestamp, whether you were mentioned or cited, competitors and the raw answer. Then compare it with the prior month. That handles reporting without pretending reporting is the work.

 

When the sheet shows you missing from 14 buyer prompts and nobody ships the pages, more frequent checks will not help. You need execution, a record of what changed, and someone answerable for the result. That is the work groas does continuously with a named strategist; I also walk through the measurement side in how to measure AI search visibility without trusting one screenshot.

 

Do not hire anyone just to count. Hire to change the count. If you cannot say which prompt or citation a tool is helping you improve, keep the cheaper sheet until you can.

 

What should a monthly report contain?

If it does not fit on one page, nobody will act on it. I used to send 14-page PPC decks with impression share by hour and auction insights by device. Clients forwarded them to nobody and changed nothing. Your AI report needs five lines:

 

  1. Mention rate: The fraction of repeated runs that named you, for each buyer prompt and engine. Say 7/10 on ChatGPT and 2/10 on Gemini, not “45% visible.”
  2. Citation rate: The same fraction for your URL in sources on those prompts. Keep it distinct from a mention.
  3. Share of voice: Your appearances beside two named competitors you actually lose deals to, not a directory of every brand an answer has ever named.
  4. What shipped: The page or citation work completed since the last report, with dates. Without this line, a change in the chart has no obvious action to inspect.
  5. What happened after: AI referrals and assisted conversions in site terms, not another rank screenshot. This is where visibility has to meet the rest of the business.

Put raw answers in an appendix if someone needs to inspect them. Cut blended visibility scores, prompt totals padded with branded questions, sentiment pies, word counts, 20-name competitor lists and charts without a before date. Some of those details can help diagnose a problem. None should crowd out the five things that tell you where to act.

 

A per-engine row beats a blended score. An average can sit still while the engine your buyers use goes to zero. The point of a monthly report is to catch that and give somebody a job to do, not to produce a defensible stack of slides.

 

How do I know a change caused a citation?

You do not know from one screenshot. I used to claim PPC wins because CPA dropped the week after I rebuilt an ad group. Then it reverted the next week when the auction normalized. AI answers make that mistake even easier: they vary between runs, and engines update without asking for permission.

 

Write down what changed and when before you take credit. Look for a sustained move across repeated samples, record model version and location, and rerun about seven days after publishing an updated page. One good answer is variance. Ten good answers after a dated change are better evidence, provided you are asking the same question under comparable conditions. They still do not give you a license to call every later lead yours.

 

My log is boring on purpose: date, page or citation changed, frozen prompt set, scores for two weeks before and two weeks after, raw answers saved. If you changed three pages at once, write down all three rather than assigning the result to your favourite. The log cannot make AI answers behave like a controlled lab test. It can stop you from turning a convenient sequence of events into a victory story.

 

Then check what happened downstream: GA4 AI channel grouping and referrals, Search Console long-query growth, and server logs to see whether the crawler saw the new page. A citation on a buyer prompt matters more when you can connect it to a shipped change and see whether anything useful followed. If the answer changed but your site data did not, keep looking. The screenshot is not the finish line.

 

No log entry, no credit. You can still be pleased when your brand appears. Just do not call it the result of your work if you cannot point to the work.

 

Which one buyer prompt, if we held the citation tomorrow, would change revenue?

That is the question I wish people asked. I used to start clients with 40 prompts because more felt thorough. It delayed the work that mattered by a month. Pick the three questions a buyer asks right before paying, score those twice as often, and fix the pages behind them first. Keep the wider list as a check, not as an excuse to postpone the obvious job.

 

Say you run a home services shop spending $20k a month on search. Your money prompts might be brutal and specific: “best installer for X near me,” “X versus Y for old houses,” and “what should X cost without the upsell?” Check those on ChatGPT and AI Overviews twice a week, 10 runs each, and write down the fractions. If a real comparison page with prices and photos goes live, put its date beside the citation rate on those prompts. A move from 2/10 to 7/10 is worth investigating; it is not proof by itself that the calls moved with it.

 

That is the habit. Count what a buyer might act on, rerun it until one lucky answer stops looking like a trend, and log what you changed with a date. Everything else risks becoming the 1am report I used to send: pretty, defensible and useless.

Frequently asked questions

What is the difference between a mention and a citation in AI search results?

A mention means your brand appears in the answer text, while a citation means your URL appears among the answer's sources. A recommendation is when the answer explicitly suggests your business, and a source appearance is when your brand shows up on a third-party page the answer cites. Mentions and recommendations are outcomes; citations and source appearances are things you can investigate and fix.

How many times should I run a prompt before I trust the AI answer about my brand?

Run each prompt many times. SparkToro found identical brand lists in the same order only about 1 in 1,000 runs, and AirOps found only about 30% of brands appear in consecutive answers. Record the percentage of repeated runs in which you appear, per buyer prompt and per engine, rather than treating a single appearance as a trend.

How do I start tracking my brand in AI answers without buying a tool?

Take 10 to 20 buyer prompts and run them in clean sessions with no history across ChatGPT, Perplexity, Gemini and AI Overviews, then score mentions, citations, recommendations and which competitors appeared. Put each prompt, engine, date and raw answer in a spreadsheet row and rerun the money prompts twice. That takes about an hour a month, costs nothing, and automating the reminder matters more than automating the running.

How should I choose which prompts to track for AI brand visibility?

Build a list of 15 to 50 buyer questions across category, comparison, alternatives, problem and branded prompts, weighted toward commercial intent, and keep branded questions separate so they cannot inflate your unbranded trend. Freeze the list once you have a baseline and keep the same prompts and engines in the same order month to month, recording mention rate, citation rate, share of voice, answer position and sentiment per row.

How can I tell whether my page update actually caused a new AI citation?

Write down what changed and when before taking credit, record the model version and location, and rerun the prompt about seven days after publishing the updated page. Look for a sustained move across repeated samples: one good answer is variance, while ten good answers after a dated change are better evidence. Then check GA4 AI referrals, Search Console long-query growth and server logs to see whether anything followed on your site.