They bought an AI visibility tracker to get cited in buyer answers. Six months later, they had a weekly chart and zero citations on the prompts that mattered.

This is a composite, not an incident from one of my accounts. It follows a sequence that is easy to recognise: buy the tracker, choose comfortable prompts, read the chart, leave the fixes unowned, face the renewal. The failure is not doing nothing about AI search. It is paying to monitor instead of paying to change the result.

The setup: a working tracker with nobody to act on it

No single business gave me this story. I built the sequence from the way monitoring tools are pitched and used, alongside a particularly clean example: a startup buys Peec AI for around 89 euros, sees zero citations, discusses what to do, then cancels three months later without making the content changes. Nobody knew how to write the content that might fix the gap.

The tool worked. The number did not move. Keep that distinction in mind as the months pass. A tracker can be accurate about a problem and still be a poor purchase for a team that has no way to act on it. That is not a software bug. It is a handoff nobody designed.

Month 0: Buying visibility when the job was earning citations

Say you run a $20k/month home services account, or a small SaaS doing $40k MRR. A buyer asks ChatGPT or Google’s AI Overviews for a recommendation, and you want your business named in the answer. That is the outcome. The team in this composite buys a tracker to see whether it happens.

The demo makes that feel like a complete plan. There is share of voice, sentiment, a competitor line, perhaps a neat view of what moved this week. But monitoring shows where a brand appears and flags gaps; it does not execute the fixes. The dashboard answers where are we missing? It does not answer who will make us worth citing there?

I understand the purchase. A visible problem feels more manageable than an invisible one. The wrong call is treating the purchase as the work rather than as a way to locate it. Before paying for the chart, the team needed to ask what would happen after it found a gap: which page would change, who could make that change, and when it would ship. Those questions are less fun in a demo because they involve people, calendars, and permission to edit a page. They are also where a measurement program becomes an operating model.

They did not ask. The tracker went live anyway.

Month 1: The first green score

The team types the prompts it knows: its brand name, its brand plus pricing, its brand against one competitor. Fifteen prompts, maybe twenty. The score comes back green.

Of course it does. Ask an engine about you by name and it will usually find you. That is not evidence that an unfamiliar buyer will find you. It is spelling.

The missing prompts are the ones that could change the business: category comparisons and problem questions from people who have never heard of the brand. The working baseline I use is 10 to 30 buyer prompts across brand questions, category and comparison questions, and pre-awareness problem questions. Hold that panel steady long enough to see a trend rather than rewriting the test whenever the answer disappoints you.

I used to make a version of this mistake with broad match. I told clients early wins on their own name meant the structure was working. I was wrong. We were measuring the part of the funnel with no competition in it. If you want to see the kind of question a brand-only list misses, read this open letter about the owner ChatGPT named your competitor instead. It is a better test of the prompt list than another reassuring brand query.

The first sign was a green score on the wrong questions. The team called it a baseline and moved on. From then on, every weekly change would be measured against a panel built to flatter the brand rather than test whether a buyer could find it.

Months 2–3: The weekly chart becomes the work

By week five, the Monday login is a ritual. Visibility goes from 12% to 14% to 11%. Someone posts the up week in Slack. Nobody posts the down week. Two months in, the report says momentum.

I would call it noise until the team shows me otherwise. LLMs can return different answers across runs, with reported accuracy swings of up to 15%. That same analysis calls for 50 to 100 runs of a prompt to estimate probability of inclusion. A weekly screenshot is one roll. It cannot carry the weight of a trend claim by itself.

The trouble grows when the team watches one prompt on one engine as if every answer were a stable result. In a study of local queries, repeat runs in Gemini showed overlapping sources about 40% of the time and the same top business about 7%, compared with about 90% repeatability for the Google local pack. Across engines, Gemini and ChatGPT cited the same domains only 8% of the time and named the same top business 4.2% of the time. Individual prompts can also show 40% to 60% monthly variation.

That does not mean measurement is hopeless. It means the unit of judgment cannot be Tuesday’s answer to this one question. Cluster the prompts by topic, keep the panel fixed, and read a 30-day rolling average at cluster level. I set out that approach in how to measure AI search visibility without trusting one screenshot.

The team makes the wrong call here: it treats movement in the report as movement in the market. That gives everyone an update to discuss without requiring anyone to touch the pages the tracker was meant to improve. Hold the test still before you call a wobble progress.

Months 4–5: The tracker finds the gap, then the work stops

By month four, the dashboard has done its job. A competitor appears on seven comparison prompts; this business does not. Reddit threads and a directory page appear where its service page is absent. The report has something useful to say.

Then nothing happens for six weeks.

The founder assumes the SEO freelancer will rewrite the pages. The freelancer assumes the tool will tell them what to do next. The tool is waiting for someone to act on the gaps it found. Monitoring can show the gap and suggest a fix, but the team still has to do the work. Assigning that work vaguely to the team is a reliable way to leave it undone.

The headline score helps everyone avoid the awkward question. A single brand visibility percentage hides which prompts and cited sources produced it. A direct recommendation is not the same thing as an appearance in a list or a passing reference. One weighted model counts those as 1.0, 0.4, and 0.2 respectively, before other adjustments. A flat mention count treats each as one.

So the team debates a two-point change in mention rate while the buyer-prompt citations it wanted remain at zero. The chart supplies numbers to discuss. It does not supply a person who can ship the missing page work. Worse, the discussion now has a familiar agenda: explain the score, compare it with last week, promise to revisit the gaps. Each meeting ends where the previous one ended.

This is the point where I want the report turned into assignments. For each weak prompt cluster, name the page that should answer it, the person who can change that page, and the ship date. Then check whether the cited-source gap changes. I point people looking for that execution side to tools that get you featured in Google AI Overviews, not just monitored. The useful distinction is not one dashboard versus another. It is measurement with an owner versus measurement without one.

Month 6: A renewal slide instead of an answer

Month six brings a renewal quote and a slide showing visibility up from 28 to 42. The account lead asks what those numbers mean for citations on buyer prompts and for AI referral traffic. Silence.

The team never kept those questions separate. Was the brand mentioned or cited? Did anyone click through? Was the relevant page crawled at all? A higher mention rate does not answer the other questions; AI referral traffic can sit flat while visibility rises. In this composite, that is exactly the distinction the renewal slide cannot resolve. The score may describe a change in what the tracker saw. It cannot stand in for the outcome the team bought the tracker to pursue.

The vendor compares its 42 with a competitor’s 38 from another tool. That does not settle anything either. Vendors use different prompt sets, engine weightings, and run frequencies; API pulls and screen captures can also return different answers. Fix the prompt set, engines, and locale, then read one series over one window. A 42 in one tool is not a 42 in another.

Six months. One subscription. Twenty-odd prompts checked weekly. Citations on the buyer prompts that decided the exercise: zero.

The team can cancel and call AI search hype, or renew at a discount and keep watching. Neither choice repairs the operating model. Tracking without an owner and a fix cadence is a paid record of inaction.

Dusty visibility dashboard beside an empty citation inbox tray

Before renewal: the checklist the chart cannot run

I would run this before buying, renewing, or building an in-house tracker. It is not a better-looking report. It is a check on whether the report can lead to shipped work. The last item is the one this composite team missed, and the earlier items cannot make up for it.

  1. Use 10–30 buyer prompts, not variations on your own name. Cover brand, category and comparison, and pre-awareness questions.
  2. Keep the prompt set, engines, and locale fixed for 30 days. Editing the panel to chase the score resets the baseline.
  3. Judge topic clusters on a 30-day rolling average. Do not promote a single weekly prompt change into a trend.
  4. Name the page that should own each weak cluster. A brand-level percentage is not a work assignment.
  5. Give one person authority to ship the page changes. Shared assumptions between founder, freelancer, and vendor are not ownership.
  6. Put a ship date against each cluster. A fix with no date remains a discussion.
  7. Define the win as buyer-prompt citations and AI referral traffic. Keep mentions, clicks, and crawl status separate.
  8. Make the same system that reports the gap accountable for changing it. This is the item most often skipped; skip it and the rest buys you another cycle of charts without fixes.

The fix: stop buying a number nobody has to move

In PPC, I would not call a search terms report a management plan if nobody could add a negative keyword. I would not call a bid report progress if nobody could change a bid. AI search deserves the same standard. The measurement can be useful, but only when it feeds work that someone owns.

That is the single rule I now apply: no prompt gets tracked unless someone is accountable for moving it. The tracker did not fail this composite team. The operating model did. It priced watching as progress and left the fixing unowned.