2022–2023: a screenshot stood in for a measurement system

A marketer typed a buyer question into ChatGPT, saw a brand name in the answer, and took a screenshot. Ten minutes later, the same question could produce a different list. That was the original AI visibility report.

Each measurement method in this timeline was built for the engine marketers thought they were measuring. Manual checks suited early answers drawn from training data. Rank-style scores borrowed the logic of fixed search results. Neither tells the whole story now that engines can assemble answers using live retrieval. If you still measure AI search as though it were a static list of names, you are reading the dashboard of a search engine that no longer exists.

Early ChatGPT had no live web access. If someone asked for the best CRM for commercial solar installers, the model answered from patterns learned during training, not a fresh inspection of vendors’ sites. A company that appeared in those answers benefited from an existing footprint across material the model had encountered. A new product could not make a quick site update and expect the answer to catch up.

That shaped the available measurement. Marketers tried five or ten versions of a buyer query, noted which brands appeared, and pasted responses into spreadsheets or decks. It was a way to inspect an answer, not a way to measure a durable position. Responses varied, and there were no outbound clicks or referral sessions to connect a mention to a business outcome.

Printed chat responses beside a highlighted spreadsheet on a cluttered desk.

The habit that stuck was worse than the spreadsheet: one screenshot began to look like a ranking. I understand the temptation. A named competitor in an answer feels concrete. But the practical reading was narrower: this particular prompt produced this particular answer. Keep the screenshot for a narrative check; do not turn it into a traffic forecast.

May 2023: SGE made an AI answer look rankable

Google introduced Search Generative Experience (SGE) in Search Labs in May 2023. Unlike an answer tucked inside a chat window, its generated snapshot appeared alongside a familiar search results page. That gave rank-tracking software something it knew how to scrape: whether the snapshot appeared and which domains it cited.

The borrowed metric was an AI position. It looked tidy in a report. The underlying surface was not. In early testing, SGE coverage swung from 84% to 15% of queries as Google changed when and how snapshots appeared. A tracker could find a citation on one check and no snapshot on the next. Calling the first result a stable rank hid the thing practitioners most needed to know: whether that answer would appear consistently at all.

This was the first big mismatch between the tool and the engine. Traditional rank tracking assumes a relatively legible place on a results page. A generated snapshot adds another question before position even matters: did the feature trigger for this query, in this check? Record the trigger and the cited source separately. A position without that context is a number with a persuasive font.

May 2024: AI Overviews arrived, but their reporting stayed mixed

Google rolled out AI Overviews to US users in May 2024. Search Console remained essential for understanding search demand and site performance, but it did not provide a separate AI Overview performance filter. Overview activity was folded into broader search reporting, making it difficult to isolate what happened inside the generated answer.

The details matter. Under Google’s reporting rules for AI features, links within an AI Overview share its position on the results page. That position does not tell you which citation looked prominent to a reader. An Overview link also needs to be scrolled into view or expanded to register an impression. A page could help inform an answer without its link producing the impression a marketer expects to see.

The workaround: sample the Overview yourself

With no dedicated filter to consult, tools used headless browsers to check keyword lists on a schedule. They recorded whether an Overview appeared and which links it showed. That was useful directional evidence, provided nobody mistook the sample for a census of customer searches.

Three limits stayed with it:

  • The test search is not the buyer’s search. Location, device, and session context can change what appears. A browser in a data center cannot represent every prospective customer.
  • A scheduled check is a moment, not a daily impression count. Checking 500 keywords once does not establish that the same Overview appeared whenever customers searched.
  • A citation is not a visit. The link may appear without a reader seeing it, clicking it, or becoming qualified pipeline.

Search Console and browser sampling therefore answered different questions. One showed aggregate performance on Google Search; the other inspected selected generated results. Use the sample to investigate the result, not to manufacture an attribution number.

Late 2024: ChatGPT Search put live retrieval into the picture

On October 31, 2024, OpenAI launched ChatGPT Search. The important change for measurement was that a search answer could draw on current web material rather than rely solely on what a model had learned during training. Search citations and outbound links made the route from answer to website more observable, though still far from complete.

OpenAI’s named user agents reflected different jobs: GPTBot for training-related crawling, OAI-SearchBot for search indexing, and ChatGPT-User for requests associated with user activity. Those distinctions matter more than a single total labeled “AI bot traffic.” Analytics could also identify some click-outs through utm_source=chatgpt.com. At last, a team could look for an actual referral session instead of treating a mention as a visit.

Prompt tracking scaled up the old screenshot

Vendors automated the manual check: run a list of prompts through an API, parse the answers, and roll the mentions into an AI Visibility Score. It saved time. It did not solve the old measurement problem, and live retrieval added another moving part.

LLM output varies across runs and contexts. A script testing “best warehouse inventory software” does not reproduce a buyer’s session, location, or follow-up questions. A change in its score might reflect changed answers to that prompt set; it does not, by itself, prove that more buyers saw your brand or that a content edit created pipeline.

Diagram contrasting automated API prompt checks with live web retrieval and crawler requests.

That does not make prompt sampling worthless. If the model keeps confusing your pricing or describing the wrong product, the answers reveal a problem worth investigating. The mistake is promoting an answer sample into a revenue metric. Once clicks and fetches became partly observable, a prompt score had to sit beside those signals, not replace them.

2025: fetch logs exposed a different part of visibility

As AI Mode and assistants made live web requests a bigger part of search, practitioners gained a useful place to look: server and edge logs. These can show whether a named crawler or user-triggered fetcher requested a product page, pricing page, or documentation. That is evidence of machine access. It is not a transcript of the answer the customer saw.

Raw logs, however, have an immediate trap. A scraper can put GPTBot or PerplexityBot in its User-Agent header. A claimed bot name is not verification. Validate requests against the available published IP information, use forward-confirmed reverse DNS where appropriate, or apply edge verification before counting them as known engine traffic. The distinction between real crawlers and spoofed user agents matters when a dashboard turns every matching string into a visibility win. If requests may be blocked before reaching your application logs, start with an audit of whether AI search can read your site.

Verification still leaves a second sorting job: separate training crawls, search indexing, and requests associated with a user’s activity. Cloudflare’s discussion of AI bot traffic illustrates why treating those categories as one pool obscures their different purposes. A verified OAI-SearchBot request supports the conclusion that a page was accessible for search indexing. A verified ChatGPT-User request points to a different kind of interaction. Neither proves that your brand was recommended, that a person saw a citation, or that anyone bought.

This is the correction to both earlier eras. A screenshot cannot establish a persistent ranking. A fetch cannot establish a finished answer. Logs tell you what reached your site; prompt samples tell you something about the answer; referrals tell you about visits. Keep those verbs separate.

Now: measure the live machine without pretending one signal is everything

A business asking “How visible are we across ChatGPT, Gemini, and Google AI Overviews?” will not get a defensible answer from one score. Each method observes a different layer, and each has a blind spot that matters in practice.

MethodWhat it can showWhat it cannot establish
Prompt samplingHow selected answers describe and cite your brandHow often real buyers receive those answers, or what caused a score to change
Google Search ConsoleAggregate Google Search impressions, clicks, and page trendsA clean split between traditional results and AI Overview activity
Verified server and edge logsWhether identified engines requested particular pagesThe wording of the final answer, a visible citation, or a human click
Referral analyticsIdentifiable visits from conversational search and what those visits did nextEvery mention or exposure that produced no click

I would not throw away the older tools. I would give each a smaller, honest job. Sample prompts to check how an answer frames your offer, not to declare a 15% rise in attributable visibility. Read Search Console for demand and page trends, not an Overview-only click count it does not provide. Verify fetches to diagnose access and request patterns, not to announce that a recommendation happened. Then inspect identifiable referrals for the outcome the other methods cannot supply: what visitors actually did after arriving.

A three-part monitoring setup for this quarter

A small team can put that division of labor into a repeatable in-house routine without paying an agency to paste screenshots into a monthly PDF:

  1. Verify and classify requests at the edge. Where verification information is available, check AI user agents against it and log validated requests for agents such as OAI-SearchBot and ChatGPT-User. Keep training, indexing, and user-associated fetches in separate groups. Review which important pages each group requests and whether your infrastructure blocks them.
  2. Track identifiable referrals against business outcomes. In GA4, isolate sessions carrying utm_source=chatgpt.com or identifiable conversational-search referrals. Follow those sessions through form fills, qualified pipeline, and closed-won revenue where your measurement allows. Do not treat the absence of a referral as proof that no answer mentioned you.
  3. Sample commercial prompts on a fixed schedule. Check twenty high-intent queries each week under consistent test conditions. Note whether your brand appears, whether the answer describes your pricing accurately, and which competitors appear beside you. Read changes as prompts for investigation, not as a causal performance report.

The order is deliberate. First confirm access, then examine answers, then connect identifiable visits to outcomes. If a page is blocked, debating its wording in a synthetic answer misses the immediate problem. If fetches occur but your offer is repeatedly misdescribed, access alone is not the win. If an answer looks flattering but sends no identifiable visits, a visibility score cannot fill in the revenue column.

That separation is also why I am skeptical of retainers built around periodic prompt screenshots. Verification, sampling, and monitoring involve plenty of mechanical work; billing more human hours for the mechanics does not make the evidence better. groas takes the opposite operating approach across paid search and earned search visibility: continuous machine execution with a named human responsible for direction and accountability, rather than a junior manager assembling another report.

The line from 2022 to now is not “screenshots were bad; logs are good.” It is that every new architecture made the previous shortcut less reliable. Early answers invited manual checks. Generated search boxes invited rank scores. Live retrieval made access and identifiable referrals worth measuring too. Measure the fetch, sample the answer, and follow the click where you can. A dashboard that collapses all three into one AI rank is still trying to report on the search engine we had before this one.

Frequently asked questions

Why did taking screenshots of ChatGPT answers work as measurement in 2022 and 2023?

Early ChatGPT had no live web access, so answers came from patterns learned during training rather than fresh inspection of vendor sites. A screenshot could show which brands a model tended to name from that training footprint, but repeating the same question later could produce a different list, so the check showed one answer rather than a stable position.

Why is an AI position score unreliable for Google's AI snapshots?

An AI position score assumes a stable place on a results page, but a generated snapshot first has to trigger for a query at all. Early SGE testing saw coverage swing from 84% to 15% of queries as Google changed when snapshots appeared, so a recorded position could come from a check where the feature showed up and be absent on the next one. The article recommends recording the trigger and the cited source separately.

Does Search Console show how my links performed inside AI Overviews?

No. Google folds AI Overview activity into broader search reporting without a separate filter. Under Google's reporting rules, links within an AI Overview share the Overview's position, and a link must be scrolled into view or expanded to register an impression, so a page can help inform an answer without producing the impression a marketer expects.

Can I use scheduled keyword checks to count AI Overview impressions?

No, sampled checks are directional evidence, not a census of customer searches. A test search runs under different location, device, and session context than a buyer's search, and one scheduled check does not establish that the same Overview appeared whenever customers searched. A recorded citation also is not a visit, since a reader may never see or click the link.

What did ChatGPT Search change about measuring AI visibility?

Launched on October 31, 2024, ChatGPT Search let answers draw on current web material instead of relying solely on training data. Search citations, outbound links, and click-outs identifiable through utm_source=chatgpt.com made the route from answer to website partly observable. OpenAI also runs distinct user agents: GPTBot for training crawling, OAI-SearchBot for search indexing, and ChatGPT-User for user-associated requests.

Is an AI Visibility Score a reliable revenue metric?

No. LLM output varies across runs and contexts, and a script testing a prompt set does not reproduce a buyer's session, location, or follow-up questions. A score change may only reflect changed answers to that prompt set, not that more buyers saw the brand or that a content edit created pipeline. Prompt sampling is better used to investigate how answers describe your offer, such as whether pricing is described accurately.

How do I verify AI crawler traffic in my server logs?

A claimed bot name is not verification, because a scraper can put GPTBot or PerplexityBot in its User-Agent header. Validate requests against published IP information, use forward-confirmed reverse DNS where appropriate, or apply edge verification before counting them as known engine traffic. Even then, a verified fetch only shows a page was accessible; it does not prove your brand was recommended or that anyone visited or bought.

What is a practical way to monitor AI search visibility in-house?

The article suggests a three-part routine: verify and classify AI user-agent requests at the edge, keeping training, indexing, and user-associated fetches separate; track identifiable referrals such as utm_source=chatgpt.com sessions through form fills, pipeline, and revenue in GA4; and sample twenty high-intent prompts weekly under consistent conditions. The order matters: confirm access first, then examine answers, then connect identifiable visits to outcomes.