

A position report can tell you where your ad appeared. It cannot tell you where your brand ranks in a ChatGPT answer, because that rank does not exist.
That is the first thing paid search teaches you to unlearn when you start measuring visibility across ChatGPT, Gemini, and Google AI Overviews. The next two are just as uncomfortable: buyers prompt these systems with context, not tidy keyword lists, and the answer to the same prompt can change from one run to the next. Put those three PPC instincts into a spreadsheet and you get numbers that look authoritative while describing very little.
I understand the appeal. Paid search has trained us to ask precise questions: Which query triggered the ad? Did it show at the top? How often? Those questions have served us well. But an AI answer is not a fixed auction result or a page of links sorted from one to ten. If you carry the dashboard over unchanged, you can spend money chasing a phantom visibility drop while missing the recommendations buyers actually receive.
Here are the six beginner beliefs I would discard first, followed by the one I am still not sure how to replace.
This is the most familiar mistake. In Google Ads, position and impression-share metrics give you a way to talk about placement. So when a founder asks, “What position do we hold when someone asks ChatGPT about our category?” a number feels like a reasonable answer. Some trackers supply one by counting where a brand appears in a list.
A bullet number is not a model ranking. An AI answer might list tools chronologically in one run, group them by operational tier in another, and alphabetize them in a third. The order tells you how that response was written, not necessarily which product the engine considers the best fit. As legacy rank-tracking methods map fixed SERP slots onto generative answers, they risk measuring formatting instead of competitive standing.
Worse, a position number can hide the sentence around your name. If the answer recommends a rival and mentions your product second only to call it an expensive legacy option, “Position 2” is not a win. It is a neat label on a bad recommendation.

Track organic mention rate and consideration share instead: how often your brand enters the candidate set across repeated runs, and whether the answer treats it as a viable choice or a footnote. If you appear in 78% of runs as a viable solution, that tells you more than assigning each appearance a rank from one to five. Read the context before you celebrate the count.
In paid search, the keyword is the unit you buy, bid on, and report against. It is tempting to paste your top twenty search terms into ChatGPT and call that an AI visibility audit. Type best crm software, record the brands, move to the next row.
The cost is a panel that reflects your account structure rather than a buyer’s problem. Someone evaluating software might include their existing stack, headcount, budget, and the headache that started the search. Those details can change the candidate set. A short head term often invites a generic list; a prompt with buying context asks the engine to make a more specific recommendation.
For a useful first pass, I would replace the keyword export with a structured panel of conversational buyer prompts. Divide thirty to fifty prompts across four intents:

Do not throw away your paid-search terms. They can suggest topics worth testing. Just do not mistake a two-word keyword for the full conversation a prospective buyer has with an answer engine.
A screenshot is seductive. You run a prompt, see your company in the second paragraph, paste it into Slack, and mark the query as won. Two days later, a colleague runs the same prompt and your brand is gone. It feels like an algorithm update. It may just be variation between runs.
Generative engines can return different answers to identical inputs. Research by Ronald Sielinski on generative search non-determinism found that a stable citation-share estimate takes far more than a single check: 40 to 50 repeated query samples on Gemini, 100 on Perplexity, and 150 or more on SearchGPT and ChatGPT. The study identified a baseline noise floor of 5 to 7 percentage points, making small day-to-day changes a poor basis for claiming that an optimization worked.
The outputs themselves change, too. An Ahrefs study of 43,000 keywords over a month found a 70% Pointwise Change Rate in Google AI Overviews. Text refreshed roughly every 2.15 days; cited sources changed 45.5% of the time, and mentioned entities changed 46% of the time. A separate variance-decomposition study of 12,933 model responses attributed 34.8% of total outcome variance to within-prompt stochastic resampling and 0.7% to brand identity.
One run is an observation, not a baseline. For an initial directional read, repeat each prompt five to ten times across different days and keep the wording fixed. If you need a stable platform-level estimate or want to credit a small change to your work, sample much more heavily. The practical cost of skipping that distinction is simple: you can spend a week “fixing” a drop that was never there.
Paid-search teams often build around Google and treat the other engines as smaller extensions of the same plan. That habit carries over easily: combine ChatGPT, Gemini, and Google AI Overviews into one AI visibility score, then watch the line move.
But these products do not draw on identical retrieval pipelines or present sources in the same way. A BrightEdge analysis of identical prompts found that ChatGPT, Google AI Overviews, and Google AI Mode disagreed on brand recommendations 61.9% of the time. Only 17% of queries returned identical brand sets; agreement on high-intent commercial queries was 23%. Data from Profound found 11% citation overlap between ChatGPT and Perplexity.
A combined score can conceal the engine where you are losing. You might appear consistently in one product and disappear from another, yet see a fairly calm average. That average is pleasant to present and hard to act on.
Keep the same prompt panel, but report each engine separately. Compare mention rates, recommendations, and cited sources within each one before rolling anything up. If a buyer-facing gap exists only on one engine, the work starts there, not in the blended number.
In a Google Ads report, an impression means your ad entered a measurable placement. In an AI answer, your name can appear for several reasons, and not all of them help you. A dashboard that counts every appearance equally skips the part a buyer actually reads.
I separate three kinds of presence:

A citation is not a recommendation. A Semrush study with Kevin Indig examined 3,981 domain appearances across 115 prompts and found that 62% of AI citations were ghost citations: the engine linked a company’s site without naming the brand in its answer. Only 13% of appearances included both an explicit brand mention and a direct source link. In that analysis, Gemini named the brand in text 83.7% of the time but provided a source link 21.4% of the time. ChatGPT linked to sources 87% of the time but named the brand in text 20.7% of the time.
If you count every linked source as buyer exposure, you can report a healthy-looking number while the answer recommends somebody else. Log the three tiers separately, and read the surrounding language. The expensive mistake is paying to improve a metric that never tells you whether you made the shortlist.
This one usually starts in a conference room. A team writes prompts around its own feature list, then feels reassured when the model names its product. Imagine a healthcare software company testing, “What is the best HIPAA-compliant pediatric patient engagement tool with bilingual two-way SMS and Epic integration?” The prompt may produce a flattering answer. It may also describe a checklist no unprimed prospect would type.
That is the cost of guessing: you can rig a visibility test without meaning to. Buyers often begin with a messy symptom or a competitor frustration, not a vendor specification. They ask why clinics are migrating away from Kareo, whether there is a cheaper alternative to Athenahealth for a two-doctor clinic, or how to reduce appointment no-shows without hiring more front-desk staff. Those prompts test whether your brand appears before the buyer knows to ask for you by name.
Start with sales calls, closed-lost notes, and support tickets. Then map the buyer prompts you lose, not just the ones tailored to make you look good. Monitoring tools such as PromptRush can test candidate panels across ChatGPT, Gemini, and Google AI Overviews and report mention rate, competitor share of voice, and external sources. That is useful diagnosis. At groas, the work does not stop at a monitoring report: the engine identifies citation gaps and rebuilds answer-friendly content, with a human strategist responsible for direction and the commercial result.

The question I want a prompt panel to answer is not “Where do we already win?” It is “Where does a buyer ask a commercially meaningful question and get sent to a competitor instead?”
I hear the appeal in this belief, too. Paid search gives you years of established reporting conventions. AI visibility does not. Waiting for a single definitive metric feels safer than assembling a panel and living with uncertainty.
But the first month does not require an enterprise suite or a bloated retainer. It requires a repeatable protocol and enough honesty to label an early read as directional. I would do it in this order:
Measurement should point to work you can do. A dashboard that tells you every month which competitor won, without helping you address the gap, is an expensive notification service. The point of this protocol is to find the missing prompts and sources, then act on them.
This is the belief I am still not sure how to replace. I spent years thinking in click IDs and conversion pixels, and I do not have a tidy, deterministic equivalent for generative search. I would be wary of anyone selling one.
In Google Ads, the audit trail can be mechanical: a user clicks an ad with a tracking parameter, lands on a page, submits a request, and the CRM links the campaign to a later sale. An AI recommendation can take a different route. A buyer asks ChatGPT to compare inventory platforms for a boutique warehouse. The answer recommends your platform and explains the fit. The buyer does not click its source link. Two weeks later, they type your URL directly or search your brand on Google and click an ad.
The conversion can then appear as direct / none or be credited to branded search. A “How did you hear about us?” field helps, but it captures only what the buyer remembers and bothers to enter. Brand-search movement and sales conversations can offer clues alongside a clean prompt panel. They are not a click-level proof of where each dollar came from.
So I would measure recommendation patterns, keep the prompt panel consistent, and watch what happens downstream without pretending the two datasets join perfectly. Generative visibility can influence a buyer who never clicks a citation. Whether attribution tools will ever show that influence with the certainty of a paid-search click log is the part I still cannot put in a spreadsheet.
Can you track a brand's position number in ChatGPT answers?
No, there is no fixed rank in an AI answer, so a bullet number only reflects how that response was formatted, not which product the engine considers the best fit. The answer may list tools chronologically, by tier, or alphabetically between runs. Track organic mention rate and consideration share instead, and read the sentence around your name before celebrating a count.
Should I test my paid search keywords in AI tools to measure visibility?
Pasting your top paid-search terms into an AI tool reflects your account structure rather than a buyer's problem, so it is not a useful audit on its own. Replace the keyword export with a structured panel of thirty to fifty conversational prompts split across category shortlists, head-to-head comparisons, competitor alternatives, and problem-led pain points. Paid-search terms can still suggest topics worth testing.
How many times should I run the same AI prompt before trusting the result?
For an initial directional read, repeat each prompt five to ten times across different days with fixed wording. Generative engines return different answers to identical inputs, with a baseline noise floor of 5 to 7 percentage points, so small day-to-day changes are a poor basis for crediting an optimization. Stable estimates require much larger samples, such as 40 to 50 runs on Gemini and 150 or more on ChatGPT and SearchGPT.
Why can't I combine ChatGPT, Gemini, and Google AI Overviews into one visibility score?
These engines draw on different retrieval pipelines and disagree often: identical prompts produced different brand recommendations 61.9% of the time across ChatGPT, Google AI Overviews, and Google AI Mode, with only 17% of queries returning identical brand sets. A combined average can look calm while your brand consistently appears in one engine and disappears from another. Report each engine separately before rolling anything up.
If an AI answer links to my website, does that count as a recommendation?
No. A citation means the answer links to your domain as a source, while a recommendation means it presents your product as a suitable choice for the user's problem. In one study of 3,981 domain appearances, 62% of AI citations were ghost citations where the site was linked without the brand being named. Log citations, mentions, and recommendations separately and read the surrounding language.
How do I find the AI prompts that actually mention my brand?
Do not write prompts around your own feature list, because that can rig a visibility test without meaning to. Start with sales calls, closed-lost notes, and support tickets to capture how buyers actually ask, including messy symptoms and competitor frustrations. The goal is to map the prompts you lose, not just the ones tailored to make your brand look good.
Do I need a perfect dashboard before I start measuring AI visibility?
No. Start with a repeatable 40-prompt core panel split across category shortlists, head-to-head comparisons, competitor alternatives, and problem-led symptoms, phrased from sales calls and support tickets rather than marketing copy. Run each engine separately with five to ten runs per prompt across different days, then audit the cited sources on high-intent prompts where competitors win. Label the early read as directional.
How do you attribute sales that came from an AI recommendation?
There is no clean deterministic equivalent of a paid-search click log. A buyer influenced by an AI recommendation may never click a citation and instead type the URL directly or search the brand later, so the conversion shows up as direct or branded search. Measure recommendation patterns with a consistent prompt panel, watch downstream signals like brand-search movement, and do not pretend the datasets join perfectly.