The standard 60-page AI visibility audit lands with 140 findings and not one fix shipped. I chose that familiar PDF for this teardown because it captures the bargain most businesses actually get: a long task list handed back to the team that bought the audit because it had no time for tasks.
You know the shape. Cover page, health score, red and yellow dots, crawl export, schema opportunities, content gap matrix, prompt screenshots, roadmap. My test for each section is the same: could a normal team take this page and put a fix live this week? Not admire the finding. Not add it to a backlog. Ship it.
The artefact: a good question buried in a big PDF
A good AI visibility audit starts with two plain questions: did the answer mention you, and did it tell the truth about you? The industry version measures that through mention rate, citation rate, share of voice, accuracy and sentiment, plus AI impressions and AI-referred leads. I would keep those measures, especially the ones you can connect to leads and revenue.
The trouble starts when the deliverable substitutes volume for decisions. Its executive summary offers a health score. A technical crawl contributes 80-plus rows. Schema gets a chapter, the content gaps get a matrix, and 30 tracked prompts get screenshots. A roadmap at the back collects the lot. The file looks thorough because it is long. Then, like too many technical audits, it sits in a shared drive because nobody validated the findings or wrote them so a developer could act.
That is the artefact I am scoring, in the order a reader encounters it. The question is not whether any of its observations are interesting. It is whether the next page makes the next action clearer.

Page one: a health score that cannot set priorities
The executive summary gives you a number like 62/100 and calls it AI readiness. It rarely explains what would have to change to make that number 70, or what either score means for pipeline. The number supplies a sense of order without telling anyone which page to fix first.
I used to tell clients scores like that were directionally useful. I was wrong. When the score has no traffic estimate or baseline behind it, it acts more like a sales device than a diagnosis. Every finding below it inherits urgency it has not earned. A missing description on a page nobody buys from can look like a worse problem than a buyer page a crawler cannot read.
Verdict: keep mention rate and citation accuracy; skip the score. If the first page cannot name the page, the failure and the person who can fix it, the first page has not earned its place at the front.
Pages 8–30: the crawl export hides the blocking issue
Next comes the technical section: 80 rows of missing alt text, long meta descriptions, H2 order and trailing slashes. The tool assigns severity by issue type, without knowing which templates carry revenue. Low-value warnings can top the list while a rendering failure on the highest-margin page sits on page four. I have seen that ordering sink a sprint. The developer clears 40 greens and yellows; the one red that mattered does not ship.
One finding in the pile can decide whether a page is available to cite at all. As of June 2026, the major AI crawlers named here do not execute JavaScript, including GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot; Gemini via Googlebot is the exception. A page whose main content appears only after client-side rendering can reach those bots as an empty shell.
My ten-minute check comes before any severity sorting: open View Source with Ctrl+U and search for the page’s main text. If you find only a root div and scripts, turn JavaScript off and reload to see what remains. That gives the team a specific page and a specific failure to investigate, rather than another colored row to file away.
Verdict: fix fetchability first, formatting last. If the main answer is absent from the page a bot receives, the other 79 rows can wait.
Pages 31–38: schema opportunities with little to ship
The schema chapter recommends Organization, FAQ, Article and Product markup to make pages machine-readable. The mechanism sounds plausible. The evidence for a citation lift is much thinner: Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 controls and found minus 4.6% on AI Overviews, plus 2.4% on AI Mode and plus 2.2% on ChatGPT. Only the decline was statistically significant. The authors concluded there was no major uplift.
I would still keep basic Organization and Article markup where it makes the site’s facts consistent. It is cheap hygiene. I would not sell it as the reason a buyer will see you cited next week. The same goes for the llms.txt page often billed as a quick win: Google says no new machine-readable files or special schema are needed for its AI features, while SE Ranking’s 300,000-domain test found llms.txt behaved as noise and most files were never fetched.
A team could add markup this week. That does not mean the audit has identified the fix worth shipping first. Verdict: keep the two lines about consistent basic markup; cut the six-page promise of visibility. Schema is hygiene, not strategy.
Pages 39–50: a content gap matrix full of the wrong gaps
The matrix is a keyword export relabeled as opportunity. It lists 50 terms for which a competitor gets cited and you do not, each paired with volume and a difficulty score. Perhaps five rows describe a missing buyer answer. The rest are what the deck calls opportunity, which I call questions your buyers never asked.
The tell is simple: the table gives a glossary term and a purchase question equal weight. Ranking for what is agentic SEO does not win the same conversation as answering which vendor to hire. Yet both appear as gaps, ready to be turned into briefs. Before a team writes either page, someone has to decide which question a buyer would actually bring to the business.
There is a second clue in the material the audit tends to treat as an afterthought. In an Ahrefs study of 75,000 brands, web mentions correlated with AI Overview visibility at 0.664, compared with 0.218 for backlinks. Concrete additions such as statistics, quotations and citations produced 30–40% gains on a test set. Those observations point to better answers and a clearer outside record, not another 45 vocabulary pages.
Verdict: keep the five unanswered buyer questions; delete the 45 glossary terms. A gap earns a place in the plan when a buyer would pay to have it answered.
Pages 51–55: prompt screenshots that stop at observation
The prompt section shows 30 buyer questions, screenshots of engine responses and your position in each. This is where the audit quietly reveals that it has no control group. Entry tracking runs from $29 a month for 15 daily prompts to $95 a month for 50 prompts across three engines, and LLM responses vary from run to run. One screenshot cannot tell you whether an apparent miss persists. Repeated runs at least give the team something sturdier than a single frame.
But the larger problem survives better tracking. The list records where you placed; it does not rewrite a weak answer, add a missing fact or publish the page that might earn a mention next time. If a prompt shows that an engine answers a buyer’s question without you, the useful deliverable is the proposed response and the page that needs work. The screenshot is evidence for that decision, not the decision itself.
There is a working method for setting up tracked buyer prompts that repeats runs instead of relying on one screenshot. Use the tracking to choose an edit, then check the answer again. Verdict: keep the buyer prompts, but require a proposed fix beside every persistent miss. Tracking without a rewrite is scorekeeping.
Pages 56–60: a roadmap that hands the work back
The last five pages compress the audit into 140 tasks labeled quick win, medium effort and strategic. There is no owner, no estimate of pipeline and no order that separates a fetching block from copy polish. Those labels describe effort vaguely; they do not tell a team what to ship on Monday.
The measurement plan underneath is also thinner than it looks. Search Console’s generative-AI report shows AI Overview and AI Mode impressions by page but no clicks, CTR or queries; pairing it with a GA4 AI-referrer channel and CRM hidden fields is how the team can follow what turned into work. Without that connection, the roadmap cannot tell which fix paid. A page can move up the task list because a tool called it severe, not because it matters to the business.
I have watched this kind of handoff die. The team bought the audit because it had no spare time. It receives 140 tasks and, sensibly, does none of them. Verdict: a roadmap without an owner and a revenue order is a second audit, not a plan. Put the first fix, its owner and its measurement on the page before adding task 140.
What survives the PDF
Three findings pass my test. First, the fetchability check: if bots cannot read the page, there is no point polishing its answer. Second, the buyer-question gap. Say you sell commercial HVAC and engines cite two competitors for restaurant grease hood cleaning cost while you have no page for it. That is not keyword filler. That is a missing sales rep. Third, the citation check: if engines already encounter consistent descriptions of a business elsewhere, the site is not making its case alone.
Notice what these have in common. Each names a piece of work: repair the page bots receive, publish the answer buyers seek, or earn a relevant outside mention. None needs a composite health score to become urgent. A useful audit could fit on two pages: one for the fetching block, missing buyer answer and off-site record; one for who ships each fix and in what order, starting with the revenue page.
That is the split groas draws between handing you chores and doing the work: track buyer questions and citations, create content, fix technical gaps, earn citations and show what moved. The audit PDF stops before execution. Keep the findings that lead to shipped pages and earned mentions. Otherwise, the auditor has diagnosed the patient and sent him home to operate on himself.
The seven checks worth running before the next audit
Forget the 140-row export for a moment. Run these checks on your five highest-margin pages first, in this order. Each should either identify a fix or clear the way for the next check; a colored dot is not an outcome.
- Confirm AI crawlers can fetch the page. Check
robots.txtand server logs for GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot; an allowlist is no help if the server returns a 403. - Confirm the main answer appears without JavaScript. View source and search for the text; a blank source is a blank page to most AI fetchers.
- Put the direct answer in the first 100 words. Answer one buyer question plainly, with a price or range when the question is commercial.
- Add one statistic, quotation or cited source per page; concrete additions produced 30–40% gains on the cited test set.
- Create the missing buyer-question page. A question for which three engines cite competitors and you have no page outranks 40 formatting fixes.
- Keep Organization and Article schema consistent; treat it as hygiene, not a citation tactic.
- Earn one off-site mention in consistent terms. Use the same name, category and claim on a source the engines already cite.
The last item is the one teams skip. It cannot be finished inside the CMS in an afternoon, so it slips beneath edits that are easier to assign. The cost is that a site can pass checks one through six and still lose the citation when another answer has an outside record behind it. That is the work the PDF leaves near the back, precisely because it requires outreach rather than another row in an export.

Overall verdict: keep the fixes, recycle the score
I would keep the fetch test, the real buyer-question gaps and the citation work. I would recycle the unexplained health score, the inflated schema chapter and the one-off prompt screenshots. They may fill pages, but they do not tell a normal team what to fix this week.
For the technical side, start with how to fix technical SEO issues that hurt AI visibility. For the missing-answer side, read why ChatGPT does not mention your brand and identify the page or mention the gap calls for. My rule now is simple: an audit that ends in tasks goes back. The one that ends in published pages and earned mentions stays.
Frequently asked questions
Is the AI readiness score in an audit useful for deciding what to fix first?
No. A score like 62/100 usually has no traffic estimate or baseline behind it and does not explain what would change it, so findings inherit an urgency they have not earned. It is better to skip the score and use measures like mention rate and citation accuracy that you can connect to leads and revenue.
Can AI crawlers like GPTBot read pages that need JavaScript to render?
No. As of June 2026, GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot do not execute JavaScript, so a page whose main content appears only after client-side rendering reaches them as an empty shell. Gemini via Googlebot is the exception.
How can I check whether an AI crawler can actually see my page's content?
Open the page source with Ctrl+U and search for the page's main text. If you find only a root div and scripts, turn JavaScript off and reload to see what remains. That gives you a specific page and failure to investigate before anything else.
Does adding schema markup improve how often AI engines cite your site?
The evidence for a citation lift is thin. Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 controls and found minus 4.6% on AI Overviews, plus 2.4% on AI Mode and plus 2.2% on ChatGPT, with only the decline statistically significant. Keep basic Organization and Article markup as cheap hygiene, but do not expect it to drive citations.
Should I create pages for every keyword gap an AI visibility audit finds?
No. Most rows in a typical gap matrix are glossary terms your buyers never asked about, and ranking for them does not win purchase conversations. Keep only the gaps where a buyer would actually pay for the answer, such as a commercial question three engines answer by citing competitors while you have no page for it.
Why isn't a screenshot of an AI answer enough to act on?
LLM responses vary from run to run, so a single screenshot cannot tell you whether an apparent miss persists. Repeated tracking runs give sturdier evidence, and every persistent miss should come with a proposed rewrite and the page that needs work. Tracking without a rewrite is just scorekeeping.
Does getting mentioned on other websites help AI engines cite your brand?
Yes. In an Ahrefs study of 75,000 brands, web mentions correlated with AI Overview visibility at 0.664, compared with 0.218 for backlinks. Use the same name, category and claim on a source the engines already cite, because a site can pass every on-page check and still lose the citation when another answer has an outside record behind it.




