web analytics
Back to blog

Prompt Research Is the New Keyword Research. Here’s My Process.

Extractable summary

Prompt research is the process of collecting the questions your buyers ask AI models, running them across engines on a fixed schedule, logging who gets cited, and turning the gaps into briefs. It replaces nothing in keyword research; it answers a different question, which is what a model says about your category when nobody from your company is in the room. This is the workflow I run: where the prompts come from, what the inventory looks like, how the test pass works, how to read the gap between the answer and what you’d have said, and how long the whole thing takes.

The average ChatGPT prompt runs about 60 words against 3.4 for a Google search, so it arrives carrying context, constraints and a decision the person is trying to make.

What most published guides skip is that the output isn’t a volume estimate. Nobody has years of prompt volume data, most individual prompts appear only once, and answers move between platforms and between reruns of the same prompt. What you get instead is a picture of what a model tells your buyer today, who it credits, and which of your positions never make it into the answer.

Where do the prompts come from?

Not from a keyword tool with question marks bolted on. That produces prompts nobody types, and the whole method collapses if the inventory is fictional.

Ranked by how much the output is worth:

  1. Recorded sales calls. The best source available and the one almost nobody uses. Twenty calls gives you the objections, the trigger language and the exact phrasing buyers use, which is what you paste into the model word for word.
  2. Support tickets and CS notes. Post-purchase pain, which tells you what the pre-purchase prompts sound like.
  3. Customer interviews, six to eight per persona, asking about the last time they evaluated a product like this. Don’t ask what content they’d like to see. People are unreliable narrators of their preferences and excellent narrators of their own history.
  4. Community threads. Reddit, Slack groups, industry forums, where you get the unfiltered question with all the constraints people won’t say on a sales call.
  5. Your own team. Sales and CS answer these questions daily and can list twenty from memory in one meeting.
  6. Search data as raw material. Search Console queries and PAA boxes give you topics, not prompts, so expand each one into the full question a person would ask with their context attached.

The test for anything entering the inventory: would a buyer type this sentence, in these words, at some point in their evaluation?

What goes in a prompt inventory?

More than the prompt. The inventory is a tracking document, so each row has the context you’ll need to read a result six months later.

FieldWhat it holds
PromptThe full question, verbatim, as a buyer would phrase it with their context included.
SourceSales call, ticket, interview, community, team, or expanded from search data.
PersonaWho asks it, and where they sit on the buying committee.
StageAwareness, consideration, or decision.
Prompt typeExplanation, comparison, or recommendation.
Our targetWhether the ideal answer cites us, mentions us, or reflects our framing without naming us.
OwnerWhich pillar and which piece is supposed to win this.
StatusCited, mentioned, absent, or misrepresented, updated each pass.

The prompt type column carries most of the strategic weight. Explanation prompts (“what is X”) produce definitions and rarely a shortlist, while comparison and recommendation prompts are where a model weighs alternatives and starts naming brands, which is the moment visibility actually matters.

30 to 60 prompts is the working range for a lean function, weighted toward comparison and recommendation, covering every persona and stage. Below 30 you’re reading noise as signal, and above 60 the monthly rerun eventually stops happening.

An inventory you rerun every month at 40 prompts beats one you built at 200 and abandoned in week six. Pick a size you can hold for a year.

How do you run the inventory across engines?

Same prompts, same order, four engines, logged in the same place every time. Consistency in the method is the only thing that makes month-to-month comparison mean anything, given how volatile the answers are.

  • ChatGPT, Claude, Perplexity and Gemini as the standing set. Cross-platform overlap is poor, with only around 11% of domains cited by both ChatGPT and Perplexity, so a single-engine read tells you nothing about the others.
  • Logged out, fresh session, memory and personalisation off. Otherwise you’re measuring what the model learned about you.
  • One prompt, one session. Follow-ups in the same thread inherit context and stop being the prompt you’re testing.
  • Log the whole answer, not just whether you appear. The text of the answer is the research output, which the next section is about.
  • Same day of the month, every month. Answers move week to week for reasons you can’t see, so a fixed cadence at least keeps the noise consistent.

For each prompt and engine, log four things: were we cited, who else was cited, what did the answer claim, and how many of the cited sources were third-party aggregators instead of vendors. That last one tells you whether the category is being narrated by review sites, which is a strategic problem and not a content one. The tracking side of this, once the inventory exists, is in how I track AI citations, and the wider measurement picture is in measuring AI visibility when analytics won’t show it.

How do you read the gap between the answer and what you’d have said?

This is the part that produces content, and it’s the step tool-led guides skip entirely, because a tool can tell you whether you were cited but can’t tell you whether the answer was any good.

Read each logged answer and mark it against four failure types:

  1. The answer is right and we’re absent. The most common result and the least interesting. The substance exists elsewhere, we’ve published nothing a model needs, standard gap, standard brief.
  2. The answer is right and we’re cited. Log what got cited and why, then work out whether it reflects the position you want to hold or some incidental fact you happened to publish two years ago.
  3. The answer is incomplete, meaning it covers the obvious considerations and misses the one that matters most in practice, which is usually the constraint that only shows up during implementation. This is the highest-value gap in the process, because you know what the consensus doesn’t and no model can reconstruct it from public sources.
  4. The answer is wrong about us, or about the category. Wrong pricing, outdated capability, a competitor’s framing applied to your product.

The question to keep in mind while reading: what would I have said if this buyer had asked me directly, and how much of that is missing? The distance between those two is your content plan. Where the distance is zero, the model already says what you’d say, so publishing it again earns you nothing, which I’ve argued at more length in why AEO won’t save a weak content strategy.

How do you convert findings into briefs?

By treating the gap as a candidate and not a commitment, and running it through the same gate as every other idea.

Each gap becomes a one-line proposition (this prompt, this persona, this missing substance) and then goes through impact, effort and alignment, the Content Prioritisation Framework I run on everything. Prompt research generates far more candidates than a lean function can make, and a research process with no kill step just produces a bigger backlog.

The briefs that survive carry three extra fields the standard template doesn’t have:

One warning from watching this go wrong. A gap list read literally produces a hundred thin pages, one per prompt. Prompts cluster, and eight related ones usually resolve into a single substantial page that answers all of them, which extracts better than eight shallow ones anyway. So cluster before you prioritise, not after, and where the gap is an information hole nobody has filled, the answer is often original research instead of another synthesis piece, which I’ve covered in why original research is the strongest AEO play.

How long does a prompt research pass take, and how often should you rerun it?

Numbers from running this solo, with nothing rounded down to make it sound easier:

StageFirst passMonthly rerun
Sourcing prompts6 to 10 hours across calls, tickets, interviews and team input30 minutes, adding new prompts as they show up
Building the inventory2 to 3 hoursNegligible
Running 40 prompts across 4 engines4 to 6 hours, since you’re reading answers and not skimming them3 to 4 hours
Reading gaps and clustering3 to 4 hours1 to 2 hours
Converting to briefs2 to 3 hours1 hour
TotalAround two working daysHalf a day

Monthly for the run and the log, quarterly for a proper review of the trend and a refresh of the inventory itself. Read the trend across three months and ignore any single month. Anything faster than monthly burns the time without adding signal, and anything slower than quarterly means you find out about a wrong answer a season late.

Now the part that gets left out of the conference talks. This doesn’t replace keyword research, it sits alongside it. Bottom-funnel search still converts, AI Overviews trigger on roughly 13% of Google queries, and walking away from the pages nearest the purchase to chase citations trades pipeline you can measure for a metric you can barely see. The distinction between the three disciplines is in SEO versus AEO versus LLMO, and the implementation errors are catalogued in the AEO mistakes playbook.

Frequently asked questions

Prompt research is the process of collecting the questions buyers ask AI models, running them across engines on a fixed schedule, logging whether you’re cited and who else is, and converting the gaps into content briefs. It differs from keyword research in both unit and output: prompts are long, contextual questions instead of short phrases, and the result is a picture of what models currently say about your category rather than a volume forecast. There’s no reliable historical volume data for prompts, so the method reads direction and pattern, never fixed positions.

Between 30 and 60 prompts works for a lean content function, weighted toward comparison and recommendation prompts over explanation prompts, and covering every persona and funnel stage. Below 30 you’re reading noise as signal, and above 60 the monthly rerun quietly stops happening, which is what kills the process. Each row should carry the prompt verbatim plus its source, persona, stage, prompt type, target outcome, owning pillar and current status.

Run the inventory monthly and review the trend quarterly, reading movement across three months instead of reacting to any single month, since AI citation rates are volatile between platforms and between reruns of the same prompt. A first pass takes roughly two working days for sourcing, inventory building, testing and brief conversion; the monthly rerun runs to about half a day. Quarterly, refresh the inventory itself by adding prompts that have come up in new sales calls and retiring ones that no longer match how buyers ask.

Prompt research doesn’t replace keyword research; it answers a different question and runs alongside it. Keyword data still governs the commercial and transactional queries nearest a purchase, which AI has disrupted least, while prompt research covers what models tell buyers during anonymous evaluation. Abandoning search work to chase AI citations trades measurable pipeline for a channel that undercounts itself badly in analytics.

How did this land?No reactions yet
Solange Rainha
Solange Rainha
Content Marketing Manager | 10+ Years B2B SaaS & AEO/LLMO