Most AEO implementation fails because teams treat it as a structural retrofit: schema, FAQ blocks, question-format headings bolted onto whatever content already exists. This article breaks down the most common AEO implementation mistakes, why the standard “AEO audit” measures the wrong things, and what the research actually says about which levers move AI citation. The fix is to invert the order most playbooks teach.
I’ll say the part that gets me in trouble at the start: AEO is real, the shift in how people find information is real, and almost every team I’ve watched start an AEO program has been doing the implementation in the wrong order.
Not because they’re lazy or stupid, but because the playbook they were handed is wrong. The dominant industry advice treats AEO as a layer you apply to finished content, a set of formatting moves you run after the writing is done: add FAQ schema, rewrite your H2s as questions, drop a TL;DR at the top, run an audit that scores how many of those boxes you’ve ticked, then wait for the citations to roll in.
They don’t roll in. The team then concludes either that AEO doesn’t work or that they need a better tool, and both conclusions are wrong.
I built an AEO/LLMO framework from scratch inside a B2B SaaS company and applied it across every content type we shipped. I wrote about the precondition for any of it working in AEO Won’t Save Your Content Strategy. Here’s What Will., which argued that structure can’t rescue a weak foundation. This piece is the sequel, and it goes one level deeper into the actual implementation: what teams get wrong once they’ve accepted that AEO matters and started doing it.
What are the most common AEO implementation mistakes?
The most common AEO implementation mistakes are treating it as a technical SEO bolt-on, over-optimizing structure while ignoring substance, and running audits that score formatting instead of the inputs that actually drive AI citation. All three come from the same root error: believing that AEO is a packaging problem when it’s a content quality problem wearing a packaging costume.
Here’s the order I see teams run, and why each step disappoints them:
- They start with the structure. Schema markup, question headings, FAQ blocks, answer-first paragraphs. This is the easy, satisfying part, because it’s mechanical and you can check it off a list. A junior marketer or a plugin can do most of it in an afternoon. So it gets done first, it gets done thoroughly, and it becomes the thing the team points to when leadership asks “are we doing AEO?”
- Then they run an audit. Usually a tool, sometimes an agency deliverable. The audit checks whether the schema validates, whether headings are question-formatted, whether there’s a summary block up top, whether the FAQPage markup is present. It spits out a score, the score goes up when you add more structure, so the team adds more structure.
- Then they wait. And the content that was thin, generic, or interchangeable before the AEO treatment is still thin, generic, and interchangeable after it. It’s just thin content with nice schema now. AI systems read the prose, and the prose didn’t change.
The whole sequence optimizes the layer that’s cheapest to fix and easiest to measure, while leaving untouched the layer that actually determines whether you get cited. That backwards sequence is the AEO implementation mistake almost nobody names, because the order feels intuitive even when it’s exactly wrong.
Why do most AEO audits miss the point?
Most AEO audits miss the point because they measure what’s machine-checkable rather than what’s machine-citable. Schema validation, heading format, and the presence of an FAQ block are all things a script can verify in milliseconds; whether your content contains a specific, attributed statistic that an AI model will reach for when answering a query is not something a generic audit script checks, because checking it requires judgment about substance.
So audits gravitate toward the checkable. You end up with a green dashboard and no citations, which is the worst possible combination, because it tells you you’re doing fine when you’re not.
I’m not saying structure is worthless. Front-loaded answers, clean headings, and self-contained sections genuinely help extraction, and I build all of them into every piece. The problem is treating that work as the finish line when it’s the warm-up. An audit that only scores structure is measuring whether your content is extractable without ever asking whether it’s worth extracting. Those are different questions, and the second one is the one that pays.
A green AEO audit on weak content is a well-formatted way to stay invisible.
If you want a test that catches this, ask of any page: strip out the schema and the FAQ block and read the raw prose. Does it have anything an AI model couldn’t already synthesize from ten other pages that say the same thing? If the answer is no, the structure isn’t your problem.
What does the research say drives AI citation?
This is where the industry consensus and the actual evidence part ways, and it’s the most useful thing in this article.
The most cited study on this is the Princeton GEO research (Aggarwal et al., presented at ACM KDD 2024), which tested content modification strategies across 10,000 queries on multiple generative engines. It’s the first peer-reviewed measurement of what moves citation rate, and almost every credible GEO playbook published since traces its numbers back to it. The findings are worth sitting with, because they contradict how most teams spend their AEO effort.
The strategies that moved the needle most were substance-side, not structure-side:
- Adding statistics improved visibility by around 41% (Princeton GEO study, KDD 2024).
- Adding direct quotations from named sources improved visibility by around 28%.
- Citing external sources improved visibility by up to 115% for lower-ranked content.
- Keyword stuffing performed worse than the baseline.
- Simply adding more words did nothing.
Read that list again with the standard playbook in mind. Statistics, quotations, and source citations are content-quality moves. They require you to have done research, to have something specific to say, to know who the credible voices in your space are and to quote them accurately. None of that is schema, none of it is heading format, none of it is the stuff the typical AEO audit scores.
The research is telling you that the levers with the biggest payoff are exactly the ones the popular playbook treats as optional. Other analyses back the same pattern: content with comparison tables earns substantially more citations than text-only equivalents, and pages that cite credible sources outperform pages that don’t. The throughline is that machines reward the same things a discerning human reader rewards, which is specificity, evidence, and a real point of view, delivered in a structure that doesn’t get in the way.
What does bad AEO look like versus good AEO?
Let me make this concrete, because “add substance” is the kind of advice that’s true and useless at the same time. The clearest way to see the AEO implementation mistakes in action is to put a weak page and a strong page side by side.
- Bad AEO looks like this: you have a blog post titled “What is customer onboarding?” You wrap it in FAQPage schema, change the H2s to questions, and add a summary box. The body says onboarding is important, that it improves retention, that companies should invest in it. Every claim is generic; there’s no data, no source, nothing a competitor’s identical post doesn’t also say. The audit scores it 95/100. It never gets cited, because when an AI model assembles an answer about onboarding, your page offers nothing it can’t get anywhere else.
- Good AEO looks like this: same topic. The piece opens with a direct, extractable answer to the question, then makes a specific claim and attributes it: a named study, a percentage, a year. It quotes a recognized practitioner by name and includes a comparison table of onboarding approaches with real trade-offs. It also has a genuine argument about which approach works for which company stage, which means there’s a reason to cite this page rather than the ten interchangeable ones. The structure is clean, but here the structure works as the delivery mechanism for everything underneath it.
The difference between those two pieces isn’t formatting. Both are well-formatted. The difference is that one of them said something and the other one performed the shape of saying something.
The content that gets cited is the content an AI model can’t reconstruct without you. Everything else is decoration on a page nobody quotes.
How does my approach to AEO differ from the industry consensus?
I invert the order. The standard playbook is: write the content, then apply AEO. Mine is: the AEO decisions are content decisions, made before and during the writing, not after. That single reordering removes most of the AEO implementation mistakes before they happen, because the substance work can’t be skipped when it’s the first step instead of the last.
In practice that means the substance levers come first. Before I worry about schema, I’m asking what specific, attributable claims this piece can make that nobody else is making. What data am I citing, and is the source credible and current? Is there original observation here, something from actual work that an AI model can’t pull from the public soup of recycled takes? That’s the part that earns the citation, so that’s the part I protect.
The structure comes second, and it’s non-negotiable but it’s also fast. Front-load the answer, write headings as the questions people actually ask, keep sections self-contained so they survive being extracted out of context, attribute every claim visibly. I documented the full version of this as a repeatable framework in Most B2B Content Is Invisible to AI Search. I Built a Framework to Change That., and I broke down the section-level tactics in Write for Humans, Structure for Machines. The structure work is real and it matters, but it’s the second move, not the first.
The other thing I do differently is measure the right outcome. A green audit score tells me my structure is clean, it tells me nothing about whether I’m being cited. For that I actually track AI citations directly, which I wrote about in How I Track Whether AI Is Citing My Content. The audit is hygiene; the citation tracking is the actual scoreboard, and confusing the two is how teams end up proud of a number that doesn’t matter.
What I’d do differently and what I’m still unsure about
I’ll be honest about the limits here, because anyone selling you certainty on AEO is selling you something.
When I first built my framework, I over-indexed on structure too. The six pillars I documented are heavy on extractability mechanics, and that reflects where my head was at the time. If I rebuilt it today, the substance levers, original data, named sourcing, genuine argument, would sit at the top of the framework rather than being implied. The structure pillars would stay, but they’d be clearly labeled as the second-order work they are.
I’m also wary of the citation numbers everyone quotes, including the ones in this article. The Princeton study is solid, but it’s from 2024, the engines have changed since, and citation rates fluctuate heavily month to month. Anyone treating a single percentage as gospel is missing that this is a moving target. The directional finding holds, substance beats structure, but the exact magnitudes are softer than the confident playbooks suggest.
What I’m confident about is the order: build the substance, then structure it cleanly, then measure whether you’re actually being cited rather than whether your schema validates. That sequence has held up across every content type I’ve applied it to, and it’s the opposite of what most teams are doing.
Frequently asked questions
The biggest AEO implementation mistake is treating AEO as a structural retrofit applied to finished content rather than a set of content-quality decisions made during writing. Teams add schema, question-format headings, and FAQ blocks to thin or generic content, run an audit that scores formatting, and then wonder why the citations never come. The structure work is necessary but it’s the second-order task; the substance, meaning specific data, named sources, and a genuine argument, is what actually earns AI citation, and it has to come first.
AEO audits don’t predict citation because they measure what’s machine-checkable rather than what’s machine-citable. A script can validate schema, confirm heading format, and detect an FAQ block in milliseconds, so audits gravitate toward those signals and produce a score that rises as you add structure. None of that checks whether your content contains the attributed statistics, expert quotations, and credible sourcing that the Princeton GEO research identified as the strongest drivers of AI visibility. A high audit score on weak content is a well-formatted way to stay invisible.
No. The Princeton GEO study (KDD 2024) found that substance-side moves like adding statistics, quotations, and source citations drove the largest visibility gains, but that doesn’t make structure worthless. Front-loaded answers, question-format headings, and self-contained sections genuinely help AI systems extract your content cleanly. The point is sequencing: structure is fast, cheap, and necessary, but it only pays off once the substance underneath it gives an AI model a reason to cite you. Teams fail when they do the structure work and skip the substance work, because clean formatting on generic content still gives an AI model no reason to cite you over ten identical pages.
