The 0–100 AI visibility score is the old Google Ads grader in a new suit. I’ve watched businesses pay to watch that number move while the citations they wanted stayed flat.

If you managed search budgets a decade ago, you remember the WordStream Google Ads Performance Grader. You connected an account, waited for a report card, and learned you had earned something like a 41 out of 100. Maybe you had fewer negative keywords than a benchmark. Maybe you had not adjusted your ad groups lately. The number looked precise, but it could not tell you whether your cost per acquisition made sense for your business or whether closed-won revenue had moved. It gave you a problem, then offered software or management services to solve it.

The trick was not that every line of the report was wrong. An account can need more negative keywords. The trick was putting a single grade above the questions that determined whether the account made money. You could raise the grade and still have a business problem. You could also have a perfectly healthy account that looked untidy to a grading rubric. Either way, the number got your attention before the context did.

Today’s AI search dashboards run a familiar script. You pay to watch a proprietary score drift from 62 to 68 across ChatGPT, Gemini, and Google AI Overviews. Meanwhile, nobody can tell you which buyer prompt you lost, who got the citation, or what changed on the page. The number moves. The work does not.

The score averages away the questions worth asking

There is no standardized industry methodology for an AI visibility score. One vendor might query thirty prompts using direct API calls; another might collect browser outputs; a third might weight one engine more heavily than another. Identical prompts can produce different citations on subsequent runs. That does not make every observation worthless. It makes a single, proprietary number a terrible substitute for a diagnosis.

Change the prompts in the panel and you can change what the number describes. Change how a mention is weighted against a citation and the same set of answers can tell a different story. This is not a complaint that measurement is impossible. It is a complaint about treating choices made inside a dashboard as though they were a verdict delivered from outside it. If I cannot see those choices, I cannot tell what the score means.

Worse, aggregated visibility indexes can blur casual brand mentions, hyperlinked citations, and placement within an answer. ChatGPT mentions your company in passing as an older provider: positive hit. Google AI Overviews links a buyer to your competitor’s pricing page instead of yours: perhaps a small dip, buried under informational queries you never cared about winning. Those are not two versions of the same outcome. One barely matters; the other might.

An analog gauge marked from 0 to 100 on a control panel with loose wires beneath it.

Practitioners comparing AI monitoring tools report sharply contradictory scores. When two dashboards looking at the same business tell opposing stories, asking which one has the nicer chart is not analysis. It is enterprise astrology with a monthly invoice.

Keep the individual observations. Throw out the magic number.

Which prompt lost, on which engine, to whom?

A composite score treats every query as an identical economic unit. Paid search managers know how absurd that is. Tell me impression share rose twelve points and my first question is: on which keywords? Did you gain commercial searches from people ready to buy, or pile up impressions on broad informational queries?

AI visibility needs the same question. Your brand might appear when someone asks ChatGPT for a definition of cloud warehousing, yet disappear when a buyer asks Gemini to compare vendor pricing. The dashboard averages both prompts into a 67 and calls it progress. A definition query is not a pricing comparison. Show me the prompt that matters, not the average that hides it.

That distinction should survive every report you send around the business. Otherwise, the person responsible for a pricing page gets the same cheerful green arrow as the person responsible for a general explainer, even when the pricing comparison is where you disappeared. The team does not need a prettier way to discuss the arrow. It needs to know which answer to inspect.

The engine matters, too. A lift from a citation to your company-history page in Perplexity does not cancel out Google AI Overviews dropping your product page in favor of software review aggregators. The composite turns two different problems into one green arrow. Try assigning that arrow to someone who has to fix the page.

And then there is the missing link: whose URL appeared instead? Engines can favor a page that answers the question more directly, including one with a useful comparison or table. Information gain is one reason a source may be worth citing. A trend line cannot show you that a competing page now has the comparison your landing page lacks. It can tell you the number fell. Helpful, if your next task is to feel bad about a number.

The cited URL gives you somewhere to start. You can read the answer, open the page it used, and ask why that page serves the prompt better than yours. Maybe the difference is obvious. Maybe it is not. Either result is more useful than an unexplained drop from 68 to 64, which gives the team nothing to work on except its feelings.

The sales demo has a familiar script

I know the presentation. Dark-mode dashboard. Animated radar sweep. A large metric labeled AI Share of Voice: 71%. The rep explains that the software checks prompts across ChatGPT, Perplexity, and Gemini so you can track your generative footprint. Then comes the part I want to see: the prompts, the citations, and the changes made because of them.

That part tends to be less cinematic.

Proprietary AI visibility scores can be hard for marketing teams to audit or tie to root causes. Look beneath an impressive index and you may find branded queries, general definitions, or prompts that do not resemble a buyer’s question. It reminds me of an old Google Ads report celebrating impression share while most of the attention came from branded searches. The report looked triumphant. The business still had to find new customers.

A presentation screen displays a rising AI Presence chart above disconnected plugs.

I am not asking a dashboard to predict every sale from every answer. I am asking it not to confuse being seen with being chosen. If the demo can show a rising line but cannot show the answers behind it, the line is decoration. Put the underlying prompts on screen. Let the buyer decide whether those are the questions worth winning.

What share of voice may actually count

Strip away the vocabulary and ask what qualifies as a win. Did your brand appear anywhere in a generated answer? Suppose a model names your product among five project management tools, then recommends two competitors for ease of use and pricing. A tracker that counts the mention gives you credit for a recommendation you lost.

A mention is not a recommendation, and a recommendation is not a citation. Counting all three as visibility is like counting every time someone says your name in a room, including when they are explaining why they bought from someone else.

Monitoring the loss does not fix it

The business-owner question is straightforward: what helps you get featured in Google AI Overviews, rather than merely monitor whether you were featured? A read-only tracker can check results and flag an omission. That has some use. But if it stops at a red arrow and a list of recommendations, the next move still belongs to someone else.

That handoff is where the sales pitch gets slippery. “We identified an opportunity” sounds close to “we improved your visibility” when both appear on the same slide. They are not close. One is a finding; the other requires somebody to change something, then look again. An alert is not an action log. If the work leaves the dashboard and vanishes into somebody else’s queue, the dashboard should not take a bow when the score recovers.

The distinction is observation versus execution. I have no objection to knowing that a citation disappeared. I object to paying for an elaborate warning system that presents the warning as the solution. Search teams have seen this before: monthly PDFs showing declining impression share, sent by people who made no bid adjustments and tested no new creative. The report was the deliverable. The account was apparently a spectator sport.

A diagnosis is not a repair

Imagine a mechanic scanning your car, confirming cylinder three is misfiring, charging you, and handing back the keys without opening the hood. That is a visibility tracker sold as an optimization tool.

Improving a page’s chance of being cited calls for work on the page and its search assets: clearer direct answers, less corporate fog, useful comparisons, structured information, and content an engine can extract. Not every change will earn a citation. At least it gives you a change to inspect instead of another color-coded chart to admire. If the tool cannot change anything, it cannot take credit for fixing anything.

A mechanic hands over an invoice while the car’s engine remains untouched.

That is the premise behind groas. Rather than sell a passive dashboard as a growth engine, groas uses purpose-built AI models to execute across paid and organic search, with a log of concrete actions and a named human strategist accountable for direction. That is the comparison I care about. Not whose index has the most decimal places, but who did the work and can show it.

What I would measure instead of a 64

When people ask how to measure visibility across ChatGPT, Gemini, and Google AI Overviews, I tell them to start by deleting the composite score. Keep a fixed panel of 20 to 30 high-intent commercial prompts: the questions prospective buyers ask when comparing solutions, prices, and vendors. Check each prompt on the engines you care about, and record what happened.

Keep the panel fixed long enough to recognize a change in the answers rather than a change in the questions you chose to ask. If you swap out the prompts every time the report looks bad, you have built your own grader. The point is not to defend a baseline. It is to see which buyer questions you are failing to answer and whether the work you did changed that.

The log needs three things:

  1. The prompt and outcome. Did the answer cite your domain with a clickable URL, mention your brand as a qualified recommendation, or omit you? Record the engine as well. Twenty prompt-level results give you twenty battles you can inspect. A blended 64 gives you a mood.
  2. The URLs cited instead. When you are absent from a valuable answer, save the sources that appeared. Compare those pages with yours. Does one answer the question directly, show a clearer feature comparison, or include a table your page lacks? That is a usable editorial brief. “Produce more content” is not.
  3. A dated log of changes. Pair page updates, content restructures, and other search work with subsequent prompt-level observations. If citation frequency rises in a week when nobody changed the site, do not award yourself a trophy; model behavior and outside sources can move without you. If you cannot point to what changed, you cannot tell whether you have a repeatable process.

None of this promises that every edit will produce a citation. It does something less glamorous and more useful: it keeps the observation, the competing source, and the attempted fix in the same conversation. When the answer changes, you have a record worth discussing. When it does not, you can stop congratulating the chart and decide what to try next.

That is less glamorous than a pulsing percentage. Good. The point is to find the lost answer, inspect the page that won it, make a change, and see what happens next. Measure the work close enough to know whether anyone did it.

Ask what they fixed last week

The next time a vendor or agency offers you an AI visibility suite, skip the tour of the radar plots. Ask one question: What did your tool or team change on a live website last week in response to a lost citation?

If the answer is an exportable audit, an alert, or recommendations for your developers to review next quarter, you have your answer. You already own software that tells you when something is broken. You do not need another bill for watching it stay broken.

A score is an observation. Execution is the product. If a platform cannot show what it changed, it is not your growth engine. It is your most expensive spectator.

Frequently asked questions

Why do AI visibility scores differ so much between tools?

There is no standardized industry methodology for an AI visibility score. One vendor might query thirty prompts using direct API calls, another might collect browser outputs, and a third might weight one engine more heavily. Identical prompts can even produce different citations on subsequent runs, so a single proprietary number is a poor substitute for a diagnosis.

What should I look at instead of one blended AI visibility number?

Look at the prompt level: which buyer prompt you lost, on which engine, and to whose URL. A composite score treats a definition query and a pricing comparison as identical, when a buyer-ready pricing question matters far more. The cited URL gives you somewhere to start by asking why that page serves the prompt better than yours.

Does monitoring my AI visibility actually improve my rankings in AI answers?

No. A read-only tracker can check results and flag an omission, but an alert is not an action log. Improving a page's chance of being cited requires work on the page itself: clearer direct answers, useful comparisons, structured information, and content an engine can extract. If the tool cannot change anything, it cannot take credit for fixing anything.

Does a brand mention in an AI answer count as a win?

Not necessarily. A model might name your product among five tools and then recommend two competitors for ease of use and pricing, and a tracker that counts the mention gives you credit for a recommendation you lost. A mention is not a recommendation, and a recommendation is not a citation, so counting all three as visibility is misleading.

How do I measure visibility across ChatGPT, Gemini, and Google AI Overviews properly?

Delete the composite score and keep a fixed panel of 20 to 30 high-intent commercial prompts that buyers actually ask. Check each prompt on the engines you care about, record whether the answer cited your domain, mentioned you, or omitted you, and save the engine along with the result. Keep the panel fixed so you see real changes in answers rather than changes in the questions you chose.

What should I do when a competitor's page gets cited instead of mine in an AI answer?

Save the URLs cited instead and compare those pages with yours. Check whether the competing page answers the question more directly, shows a clearer feature comparison, or includes a table your page lacks. That comparison gives you a usable editorial brief, whereas 'produce more content' does not.

How can I tell if an AI visibility vendor is actually doing work or just selling reports?

Ask what their tool or team changed on a live website last week in response to a lost citation. If the answer is an exportable audit, an alert, or recommendations for your developers to review next quarter, they are only observing. A score is an observation; execution is the product.