web analytics
Back to blog

How to Audit Your Content Library for AI Readiness (Without Losing Your Mind)

Extractable summary

An AI readiness audit is a page-by-page assessment of how well your existing content can be parsed, extracted, and cited by AI systems like ChatGPT, Claude, and Perplexity. This is the methodology I use to score a library against five checks, separate the pages worth fixing from the ones worth killing, and scope the work so it takes days rather than months. It turns the AEO/LLMO framework into a repeatable operational process.

You inherited a content library, two hundred pages, maybe four hundred, built up over years by people who left, optimised for a version of search that’s quietly disappearing, and now someone has read a LinkedIn post about AI search and wants to know if the library is “AI-ready,” and the answer is that nobody knows, because nobody has looked at it through that lens.

An AI readiness audit is how you look, it’s a scored assessment of whether your existing content can be lifted and cited by an AI model, applied page by page, fast enough to get through a real library without it becoming a quarter-long project. The output is a prioritised list: what to fix, what to kill, what to leave alone. I’ve run this on libraries I built myself and on ones I walked into cold, and the version below survived contact with both.

Discovery is moving off the blue link, which is what makes this urgent: around 58.5% of US Google searches already end without a click (Semrush, 2025), and Gartner predicted a 25% drop in traditional search volume by 2026 (Gartner, 2024). If a meaningful share of your buyers now ask an AI model before they ask Google, a library that can’t be cited is one that’s slowly going dark, whatever its traffic graph says.

What is an AI readiness audit, and why bother running one?

An AI readiness audit assesses how easily an AI system can find a clean, self-contained answer on a page and attribute it back to you; it checks two things at once: whether the page says something worth citing, and whether it’s built so a machine can extract that something without reading the whole thing.

Those are separate problems, and most teams only audit the second one: they run a tool that checks for schema, FAQ blocks, and heading format, get a green score, and assume the work is done. I went deep on why that’s backwards in The AEO Hype Is Real, But Most Teams Run the Playbook Backwards, but the short version is that a green audit on a thin page is a well-formatted way to stay invisible.

So the audit has to grade both: a page that makes a specific, attributed claim but buries it under three paragraphs of throat-clearing scores badly on extraction, and a page with perfect structure that only repeats what ten other pages already say scores badly on substance; what you want clears both bars, and the audit exists to tell you how many of those you have.

The volume problem predates AI: Forrester has reported for years that 60 to 70% of B2B marketing content goes unused (Forrester). An AI readiness audit doubles as a forcing function to confront that pile, because a lot of what you find won’t be worth optimising for anything, AI included.

The five checks: AI-ready versus AI-invisible

Every page gets scored against the same five checks, each worth 0, 1, or 2 points, for a page total out of 10. The checks come straight out of the AEO/LLMO Content Optimisation Framework, reordered so substance comes before structure, because that’s the order the research supports and the order teams get wrong.

CheckAI-invisible (0)Partial (1)AI-ready (2)
1. Front-loaded answerThe core question isn’t answered until halfway down, or never directly.An answer exists but is wrapped in setup or hedging.The first 40 to 75 words after the heading directly answer the question, with the qualifying context right behind it.
2. Citable substanceGeneric claims any competitor’s page also makes; nothing a model can’t get elsewhere.Some specifics, but mostly consensus restated.An original position, attributed data, or a named source a model would have to come here to get.
3. Standalone sectionsSections only make sense after reading the ones above them.Some self-contained, some dependent.Each section answers one question and survives being pulled out of context entirely.
4. Attributed evidenceOrphan stats with no source; comparisons buried in prose.Some sourcing, no structured evidence.Every figure linked to its source; genuine comparisons presented as tables.
5. Retrievable structureVague headings, definitions buried in narrative, JavaScript-rendered body.Mixed; some question headings, some buried terms.Question-format headings, key terms defined at first use, clean pre-rendered HTML.

The check that does the heaviest lifting is the second one, citable substance, and it’s the one a structural audit can’t see. The Princeton GEO study (Aggarwal et al., presented at ACM KDD 2024, tested across 10,000 queries) found the levers that move citation are substance-side: adding attributed statistics lifted visibility by around 41%, direct quotes from named sources by around 28%, and citing external sources by up to 115% for lower-ranked content, while adding more words did nothing and keyword stuffing scored below baseline (Princeton GEO study). Treat those magnitudes as directional, since the engines have changed since 2024, but the ordering has held up across every library I’ve applied it to.

How do you score a page in two minutes?

You read the first screen, skim the headings, and check three things, because that’s where most of the signal lives: an analysis of LLM citations found that around 44% are pulled from the first 30% of a page (Virayo, 2025), so if the opening is weak, the rest of the page barely gets a vote.

  1. Open the page and ask whether the first paragraph answers the question the title implies or just warms up.
  2. Scan the H2s next: are they questions a buyer would actually type, or label-nouns like “Overview” and “Benefits”?
  3. Last, look for one real number with a source attached.

Those three reads cover checks one, five, and four, and they take about ninety seconds. Substance and standalone sections take another thirty seconds of judgment, and then you write the score down and move on. Resist the urge to fix anything while you score; auditing and fixing are different jobs, and mixing them is the fastest way to turn a one-week audit into a two-month one.

Once a page has a score, it lands in one of three buckets:

  1. AI-ready (8 to 10). Leave it. It’s doing its job. Note it as a model for the rest of the library.
  2. Salvageable (4 to 7). This is where your effort goes. The substance is mostly there and the fixes are mechanical.
  3. AI-invisible (0 to 3). Either rewrite from scratch or kill it. Optimising a page that says nothing is polishing a thing nobody will cite.

The salvageable bucket is the whole game: a page scoring 6 usually has the substance and just needs its answer pulled to the top and its evidence sourced, which is an hour of work that can move it to a 9, but a page scoring 2 needs to be rebuilt or removed, and rebuilding is a new-content decision, not an audit fix.

What to fix first: the 80/20 of AI readiness

Fix the salvageable pages that already get traffic or map to a real buying decision, in that order, and ignore almost everything else for now. The audit will hand you a list far longer than you can action, so the prioritisation call is what keeps the project from collapsing under its own weight.

I score the fix list on two axes: how close the page is to AI-ready (its audit score) and how much it matters to the business (traffic, buyer intent, or strategic relevance to what you’re selling now).

High business valueLow business value
Salvageable (4 to 7)Fix first. Best return on effort in the whole library.Fix later, or only if it’s a quick win.
AI-invisible (0 to 3) or AI-ready (8 to 10)If AI-ready, leave it; if invisible, schedule a rewrite as new content.Leave alone, or kill if it’s actively misleading.

This is the same logic I use for deciding what to make in the first place, which I wrote up as a full content prioritisation framework. The audit version is simpler because the content already exists, so effort is mostly known and the only question is impact: a high-traffic page sitting at a 5 is the single best use of your time, because the traffic proves it has authority worth making citable, and a one-hour fix can convert that.

The trap is the inverse: a stakeholder’s favourite page that scores well but matters to nobody, or a pristine pillar piece on a topic you’ve since deprioritised. Score it, note it, leave it; effort spent making an irrelevant page AI-ready is effort stolen from a page that could have earned citations on something you actually sell.

What are the structural failures I see most often?

The same handful of problems account for most low scores, and once you’ve seen them you can spot them at a glance.

  • The buried answer. The page knows the answer and makes you wait for it, with a 250-word introduction about “the changing world of X” sitting between the heading and the first useful sentence. Models retrieve at the heading level and favour passages that answer in one self-contained block, so the answer that arrives in paragraph four might as well not exist. The fix is brutal and fast: cut the introduction, promote the answer.
  • The orphan statistic. A page claims “73% of buyers do X” with no link, no study, no year. A human skims past it; a model has no reason to cite an unverifiable number when five other pages attribute the same figure.
  • Sections that can’t stand alone. A section that opens with “As mentioned above” or relies on a definition introduced three screens earlier can’t be lifted cleanly. The fix is to make every section restate just enough to work alone, which also happens to make the page better for skim-readers.
  • The comparison trapped in prose. The page compares three approaches across four dimensions, and it does it in flowing paragraphs. That reads fine, but it hands a model nothing structured to extract. Pages with tables get cited far more than prose equivalents of the same data; one 2025 citation analysis put tables at around 4.2 times the citation rate of prose, and answer-sized passages of 40 to 75 words at around 3.1 times longer ones (Kime, 2025).
  • The page that says nothing. No failure of structure can save this one: it’s well-written, it’s on-topic, and it contains not one claim, number, or position a model couldn’t assemble from the rest of the open web. This is the page you stop optimising and start questioning, and it’s why substance is check two and not check five.

How do you scope an AI readiness audit so it doesn’t eat six months?

You sample instead of auditing everything, you time-box the pass, and you batch the fixes. A four-hundred-page library audited page by page at full depth is a project that never ships, and a project that never ships produces no citations, so the discipline here is as important as the scoring itself.

  1. Sample the long tail, audit the head in full. Score every page that gets real traffic or maps to a current buying decision, which is usually a fraction of the library. For the long tail, score a representative sample of fifteen or twenty pages; the failures cluster, so the sample tells you what the rest looks like without you having to open all of it. I leaned hard on this kind of triage when I rebuilt a company’s content function in 90 days, where there was no time to treat every page as precious.
  2. Time-box the scoring pass. Give yourself a fixed window, a day or two, to score the head and the sample, and accept that some scores will be rough. A rough score on every important page beats a perfect score on the first forty pages and nothing after.
  3. Batch the fixes by failure type. Once you’ve scored, you’ll see the same problems repeat, so fix them in batches: one session promoting buried answers across twenty pages, one session sourcing orphan stats, one session building tables. Batching is faster than fixing one page end to end, and it’s far less mind-numbing, which is important when you’re staring down a list this long.

A content audit for AI readiness done this way is a week of work for a library that would take a quarter to do exhaustively, and the week gets you most of the value. This is the operational discipline that turns a strategic framework into something a one-person content function can actually run, which is the whole argument behind treating content as a system rather than a team.

What I’d do differently

The first time I ran an audit like this, I scored everything, and it was a waste. I spent two days on the long tail that I should have spent fixing the ten pages that mattered, and the long-tail scores told me nothing the sample wouldn’t have. Sampling is the right method from the start; full enumeration feels more rigorous, and on a long tail it mostly buys you a slow confirmation of what the sample already told you.

I’d also be more honest, earlier, about the measurement gap. An AI readiness audit scores how citable a page should be. It does not tell you whether the page is actually being cited, and those can diverge, because citation is noisy and platform-specific. Only around 11% of domains are cited by both ChatGPT and Perplexity (LLMPulse, 2025), and citation rates drift month to month, so a page that scores a 9 might still underperform on a given engine for reasons the audit can’t see. The audit gets you to “structurally and substantively citable,” but proving you’re cited needs separate tracking, which I covered in the metrics that actually tell you if your content is working.

And I’d push back harder on the instinct to optimise the whole library, because it usually comes from anxiety rather than strategy. Rand Fishkin, who has spent two decades watching this shift, puts it bluntly: “Traffic is a vanity metric” (Vende Digital interview, 2025). The same logic applies to AI readiness: chasing it across every page is a vanity goal of its own, when the job is getting the pages that matter cited and stopping there. An audit that ends with “fix these eight pages” is more useful than one that ends with “fix everything,” and I’d trust that smaller list sooner than I did.

The AEO foundations behind all of this, and why substance has to lead, are in AEO Won’t Save Your Content Strategy. The audit is how you find out where your library stands against those foundations before you spend an hour optimising the wrong thing.

Frequently asked questions

An AI readiness audit is a page-by-page assessment of how well an existing content library can be parsed, extracted, and cited by AI systems such as ChatGPT, Claude, Perplexity, and Google’s AI features. It scores each page on two dimensions: whether the page contains substance worth citing, and whether it’s structured so a model can extract that substance without reading the whole thing. The output is a prioritised list of what to fix, what to rewrite, and what to leave alone.

A content audit for AI readiness should take about a week for a typical B2B library, not a quarter. The way to keep it short is to score every high-traffic or high-intent page in full but only sample fifteen to twenty pages from the long tail, since structural failures cluster and the sample is representative. Time-box the scoring pass to a day or two and batch the fixes by failure type rather than fixing each page end to end.

A traditional SEO content audit checks for keyword targeting, traffic, backlinks, and technical health, all oriented around ranking and click-through. An AI readiness audit checks whether content can be cited in a zero-click, answer-engine context, which means it grades front-loaded answers, citable substance, standalone sections, and attributed evidence. The two overlap on structure but diverge on intent: SEO optimises for the click, AI readiness optimises for being the cited answer when there is no click.

Yes, and this is the most common false positive. A page can have perfect schema, question-format headings, and an FAQ block, score green on any automated tool, and still never get cited because it says nothing a model can’t get from ten other pages. Structure is necessary but not sufficient; substance, meaning an original position, attributed data, or a named source, is what actually earns the citation. That’s why a proper AI readiness audit grades substance first and treats the structural checks as the delivery layer.

How did this land?No reactions yet
Solange Rainha
Solange Rainha
Content Marketing Manager | 10+ Years B2B SaaS & AEO/LLMO