web analytics
Framework · Score it · Ship the fix list

AI Readiness Audit Framework.

A practitioner’s framework for scoring an existing content library on whether AI systems can parse, extract, and cite it. Read each page against five checks, triage the list by business value, then scope the pass so it finishes in a week instead of a quarter.

Score one page · two minutes, not twenty
Front-loaded answer Structure
Do the first 40 to 75 words after the heading answer the question directly?
Citable substance Substance
An original position, attributed data, or a named source a model would have to come here to get?
Standalone sections Structure
Does each section survive being pulled out of context?
Attributed evidence Substance
Every figure linked to its source, comparisons built as tables?
Retrievable structure Structure
Question-format headings, terms defined at first use, clean pre-rendered HTML?
Page score
0 / 10
0 of 5 scored

Score each check 0, 1, or 2. Two of them grade substance and three grade structure, and the substance pair does the heavy lifting.

Built by Solange Rainha

I built this AI readiness audit as the operational layer under my AEO/LLMO Content Optimisation Framework, the part that tells you where a library actually stands before you spend an hour optimising the wrong page. I’ve run it on libraries I built myself and on ones I walked into cold, four hundred pages deep, written by people who’d long since left and optimised for a version of search that’s quietly disappearing. It runs on five checks, scored fast, applied page by page, and it ends with a prioritised list rather than a forensic report nobody reads.

The fundamentals

Two halves, one verdict.

Most audit advice hands you a checklist and leaves the genuinely hard part, deciding what to do with four hundred scores, entirely to you. This AI readiness audit has two halves because the score is only useful if you can turn it into a fix list that ships.

Scoring

What does the page actually do?

A repeatable read on every page against five checks, scored 0, 1, or 2 each for a page total out of 10. The goal is a fast directional call on whether a model can find a clean answer on the page and attribute it back to you, not a precise grade you’ll end up arguing about.

Triage

Which pages earn the fix?

A library of scores means nothing until you decide where the week goes. The second half crosses each page’s score with its business value, sorts the fixable from the dead, and scopes the pass so a four-hundred-page library takes days. This is the half that keeps the audit from collapsing under its own weight.

The operating logic

Score, triage, scope, in that order.

The three moves don’t carry equal weight, and the order you run them in changes the answer. Score without triaging and you’ve built a spreadsheet nobody acts on. Skip the scoping and the fix list runs longer than your year. Run them in sequence and the audit stays a week of work instead of a standing project.

01

Score: read the page, write the number

Read the first screen, skim the headings, check three things, move on. The point is a directional read in two minutes, not a forensic one in twenty. Resist the urge to fix anything while you score, because auditing and fixing are different jobs and mixing them is how a one-week audit becomes a two-month one.

02

Triage: cross the score with what the page is worth

A score on its own doesn’t tell you where to spend the week. Cross it with business value, meaning traffic, buyer intent, and strategic relevance to what you sell now, and only one combination earns your time: a salvageable page that matters.

03

Scope: ship before it eats six months

Sample the long tail instead of enumerating it, time-box the pass, batch the fixes by failure type. A library audited exhaustively is a project that never ships, and a project that never ships earns no citations.

Layer 1 · The scoring instrument

The five checks: AI-ready versus AI-invisible.

These are the five reads that produce the page score. Each is worth 0, 1, or 2, for a page total out of 10. Two of them grade substance and three grade structure, and the substance pair does the heavy lifting, because it’s the part a structural tool can’t see and the part that actually earns a citation.

CheckAI-invisible (0)Partial (1)AI-ready (2)
1. Front-loaded answerThe core question isn’t answered until halfway down, or never directlyAn answer exists but is wrapped in setup or hedgingThe first 40 to 75 words after the heading directly answer the question, with the qualifying context right behind it
2. Citable substanceGeneric claims any competitor’s page also makes; nothing a model can’t get elsewhereSome specifics, but mostly consensus restatedAn original position, attributed data, or a named source a model would have to come here to get
3. Standalone sectionsSections only make sense after reading the ones above themSome self-contained, some dependentEach section answers one question and survives being pulled out of context
4. Attributed evidenceOrphan stats with no source; comparisons buried in proseSome sourcing, no structured evidenceEvery figure linked to its source; genuine comparisons presented as tables
5. Retrievable structureVague headings, definitions buried in narrative, JavaScript-rendered bodyMixed; some question headings, some buried termsQuestion-format headings, key terms defined at first use, clean pre-rendered HTML

Check 5 is the only check where a deliberate deduction can be the right call; the note under check 05 below explains when.

The check that carries the audit is the second one, citable substance, and it’s the one an automated tool scores green while missing entirely. The Princeton GEO study (Aggarwal et al., presented at ACM KDD 2024, tested across 10,000 queries) found the levers that move citation are substance-side: attributed statistics lifted visibility by around 41%, direct quotes from named sources by around 28%, and citing external sources by up to 115% for lower-ranked content, while adding more words did nothing (Princeton GEO study). Treat the magnitudes as directional, since the engines have moved on since 2024, but the ordering has held up on every library I’ve scored.

01 · Structure

Front-loaded answer

The first 40 to 75 words after the heading answer the question the heading implies, with the qualifying context right behind. Around 44% of LLM citations are pulled from the first 30% of a page (Virayo, 2025), so an answer that arrives in paragraph four barely gets a vote.

02 · Substance

Citable substance

The page holds an original position, attributed data, or a named source a model would have to come here to get. If the prose only restates what ten other pages already say, no amount of formatting earns the citation. This is the check a structural audit is blind to, and the reason the framework grades substance before structure.

03 · Structure

Standalone sections

Every section answers one question and survives being lifted out of context. A section that opens with “as mentioned above” or leans on a definition three screens up can’t be extracted cleanly, because the model grabs it in isolation and gets a fragment.

04 · Substance

Attributed evidence

Every figure carries a visible source, and genuine comparisons are presented as tables instead of buried in prose. An orphan stat is a liability: a model has no reason to cite an unsourced “73% of buyers” when five other pages attribute the same number.

05 · Structure

Retrievable structure

Question-format headings that match what buyers actually type, key terms defined at first use, clean pre-rendered HTML a crawler can read. This is the fastest layer to fix and the one teams over-index on, which is why it sits last. Question headings also flatten a distinctive voice, so where that voice earns the citation, take the deduction on purpose.

The evidence

What the research says about where citations come from.

The case for grading substance first, and for front-loading the answer, rests on what actually moves AI citation. The numbers are directional and engine-specific, but the direction has been consistent.

~41%

visibility lift from adding attributed statistics to content

Princeton GEO study, KDD 2024
up to 115%

citation lift from citing external sources, for lower-ranked content

Princeton GEO study, KDD 2024
~44%

of LLM citations pulled from the first 30% of a page

Virayo, 2025
~4.2x

citation rate of tables over prose presenting the same data, with answer-sized passages of 40 to 75 words cited around 3.1x more than longer ones

Kime, 2025

Only around 11% of domains are cited by both ChatGPT and Perplexity, a reminder that citation is noisy and platform-specific (LLMPulse, 2025).

Layer 2 · The delivery layer

The triage layer: turning scores into a fix list.

Scoring gives you a number per page. This layer is how that number becomes a decision that survives a four-hundred-page library and a backlog that wants to become a to-do list nobody finishes.

01 · Triage

Sort every page into one of three buckets

AI-ready (8 to 10) gets left alone and noted as a model for the rest. Salvageable (4 to 7) is where your effort goes. AI-invisible (0 to 3) gets rewritten from scratch or killed, because optimising a page that says nothing is polishing a thing nobody will cite.

02 · Triage

The salvageable bucket is the whole game

A page scoring 6 usually has the substance and just needs its answer pulled to the top and its evidence sourced, an hour of work that can move it to a 9. A page scoring 2 needs rebuilding, which is a new-content decision rather than an audit fix. Spend your week in the 4-to-7 band.

03 · Triage Decide

Cross the score with business value

Score every page on two axes: how close it is to AI-ready, and how much it matters, meaning traffic, buyer intent, and relevance to what you sell now. Only one quadrant earns your week. On a small site the page read right before a decision gets low, deliberate traffic, so audit the conversion set whatever its numbers say.

04 · Scope

Sample the long tail, audit the head in full

Score every page that gets real traffic or maps to a current buying decision, which is usually a fraction of the library. For the long tail, score a representative sample of fifteen or twenty; the failures cluster, so the sample tells you what the rest looks like without you opening all of it. I leaned on this kind of triage when I rebuilt a content function in 90 days, where there was no time to treat every page as precious.

05 · Scope

Time-box the pass and batch the fixes

Give yourself a day or two to score the head and the sample, and accept that some numbers will be rough. Then fix in batches by failure type, one session promoting buried answers, one session sourcing orphan stats, one session building tables, because batching is faster and far less mind-numbing than fixing each page end to end.

High business valueLow business value
Salvageable (4 to 7)Fix first. Best return on effort in the whole libraryFix later, or only if it’s a quick win
AI-invisible (0 to 3) or AI-ready (8 to 10)If AI-ready, leave it; if invisible, schedule a rewrite as new contentLeave alone, or kill if it’s actively misleading
Applied

The audit blueprint.

How the two layers run on a single page, start to finish.

01

Read the first screen

Does the opening answer the question the title implies, or warm up? That one read covers the front-loaded-answer check and most of the substance one.

02

Score the five checks

Front-loaded answer, citable substance, standalone sections, attributed evidence, retrievable structure. Zero, one, or two each, total out of ten.

03

Bucket it

AI-ready, salvageable, or AI-invisible. Write the number down and move on without fixing anything.

04

Cross with business value

A salvageable page that matters goes to the top of the fix list. A pristine page nobody needs gets noted and left.

05

Batch the fix

Group the salvageable-and-valuable pages by failure type and clear them in passes, not one page at a time.

06

Track whether it worked

The audit scores how citable a page should be. Whether it’s actually cited needs separate tracking, because the two diverge.

Workflow

How it fits into content operations.

The audit sits in front of any AI-readiness work, before a single page gets rewritten, and it leaves a prioritised list you can defend later.

1

Inventory

Pull the full library and tag each page with its traffic and its relevance to what you sell now.

2

Score

Five checks per page, head in full, long tail by sample. A day or two, not a fortnight.

3

Bucket

AI-ready, salvageable, or AI-invisible, with the number attached.

4

Prioritise

Cross score with business value; the salvageable-and-valuable quadrant is the week’s work.

5

Batch-fix

Group by failure type, promote answers, source stats, build tables, in passes.

6

Measure

Track actual AI citations on the fixed pages. That’s the scoreboard, not audit completeness.

This is the operational discipline that turns a strategic framework into something a one-person content function can actually run, which is the whole argument behind treating content as a system rather than a team.

Made concrete

What an AI-ready page looks like versus an invisible one.

Both can be well-written and on-topic. Only one gives a model a reason to pull from it.

AI-invisible: structure with nothing under it

  • A 250-word warm-up before the first useful sentence
  • Generic claims any competitor’s page also makes
  • Orphan stats with no source, year, or link
  • A comparison trapped in flowing paragraphs
  • Sections that only make sense read top to bottom

AI-ready: substance the structure delivers

  • The answer in the first 40 to 75 words, qualifier right behind it
  • An original position or attributed number a model has to come here to get
  • Every figure linked to its source
  • The comparison built as a table
  • Each section standing on its own when lifted out
The usual suspects

The five failures that tank a page score.

The same handful of problems account for most low scores. Once you’ve seen them you spot them at a glance, which is what makes the audit fast, because you stop re-deriving the diagnosis and start pattern-matching.

01

The buried answer

The page knows the answer and makes you wait for it, a long introduction sitting between the heading and the first useful sentence. Models retrieve at the heading level and favour a self-contained answer block, so the one that arrives in paragraph four might as well not exist. The fix is brutal and fast: cut the introduction, promote the answer.

02

The orphan statistic

A claim like “73% of buyers do X” with no link, no study, no year. A human skims past it; a model has no reason to cite an unverifiable number when five other pages source the same figure. Every external stat needs a visible source, and one you can’t source gets cut rather than left dangling.

03

Sections that can’t stand alone

A section that opens with “as mentioned above” or relies on a definition introduced three screens earlier can’t be lifted cleanly. The fix is to make every section restate just enough to work alone, which happens to help skim-readers too.

04

The comparison trapped in prose

Three approaches across four dimensions, written out in paragraphs. It reads fine and hands a model nothing structured to extract. Where a page has a comparison buried in sentences, building the table is often the single highest-value edit on it.

05

The page that says nothing

Well-written, on-topic, and holding not one claim, number, or position a model couldn’t assemble from the rest of the open web. No structural fix saves this one, and it’s why substance is check two rather than check five.

The failure modes

Where the audit breaks down.

The framework is simple. The ways people break it are predictable, and I’ve made most of these myself.

01

Scoring everything instead of sampling

The first time I ran an audit like this, I scored every page, and it was a waste. I spent two days on a long tail that told me nothing the sample wouldn’t have, time I should have spent fixing the ten pages that mattered.

The reframe

Audit the head in full, sample fifteen to twenty pages from the tail. Full enumeration feels more rigorous and mostly buys you a slow confirmation of what the sample already said.

02

Fixing while you score

The pull to fix a buried answer the moment you spot it is strong, and giving in to it is how a one-week audit becomes a two-month one. Every page you stop to repair is a page you’re not scoring, and the scoring pass is the thing that produces the prioritised list.

The reframe

Score the whole pass first, fix nothing. Batching the fixes afterwards is faster than interleaving them, and it keeps the scoring read fast and consistent.

03

Trusting a green structural score

An automated tool checks schema, heading format, and the presence of an FAQ block, all in milliseconds, and none of it sees whether the page holds an attributed number or a named source. A green score on a thin page is a well-formatted way to stay invisible.

The reframe

Grade substance first and treat the structural checks as the delivery layer. The tool measures hygiene; citable substance is what gets cited.

04

Optimising the whole library

The instinct to make every page AI-ready usually comes from anxiety rather than strategy. Rand Fishkin, who has spent two decades watching search shift, puts the broader version bluntly: “Traffic is a vanity metric” (Vende Digital interview, 2025). Chasing AI readiness across every page is a vanity goal of the same kind.

The reframe

Fix the pages that matter and stop. An audit that ends with “fix these eight pages” is worth more than one that ends with “fix everything,” and the smaller list is the one you’ll actually action.

05

Confusing “should be cited” with “is cited”

The audit scores how citable a page should be, structurally and substantively. It does not tell you whether the page is actually cited, and the two diverge, because citation is noisy and platform-specific: only around 11% of domains are cited by both ChatGPT and Perplexity (LLMPulse, 2025).

The reframe

Treat the audit as the diagnostic, not the scoreboard. Proving you’re cited needs separate citation tracking, run after the fixes land.

Honest about the limits

This audit is built for a solo or lean content function, where there’s no team to absorb a slow exhaustive pass and the only acceptable output is a short fix list. On a bigger team with spare capacity you can afford more depth, and some of the ruthlessness softens. The scoring is also deliberately directional, a two-minute read producing a rough number, so it’ll feel too loose for anyone who wants a formula that spits out a ranked queue. That looseness is the feature, because the moment scoring a page takes longer than fixing it, you’ve rebuilt the bureaucracy this was meant to kill. What holds across every library I’ve run it on is the order: score the page, triage by value, scope so it ships, then measure whether the fixes actually earned the citation. The deeper case for why substance has to lead all of this is in the AEO/LLMO Content Optimisation Framework, and the logic for deciding what to make in the first place is in the Content Prioritisation Framework.

Where this came from

Built on real libraries.

If you’re running something similar and it breaks somewhere mine doesn’t, I want to hear about it.