Service Support — AI Visibility
Our Prompt Research Process, Step by Step
Our prompt research process for AI visibility (GEO/AEO): how we find, score, and prioritise the exact prompts buyers ask AI engines, step by step.

Our prompt research process is the first thing we do on every AI visibility engagement, before touching a single page: build a list of the exact questions real buyers type into ChatGPT, Perplexity, Gemini, and AI Overviews, run every one manually to see what the engines currently say, then score and cluster the gaps into a prioritised list. Nothing gets written or restructured until this list exists — it’s what tells us which prompts are worth chasing.
Key takeaway
- Six steps: pull real buyer language, draft the raw list, run every prompt manually across engines, score by intent and value, cluster by topic, then hand off a prioritised list.
- We never start with keyword-tool exports — search keywords and AI prompts overlap but aren’t the same thing, and treating them as identical is a common mistake.
- The output isn’t a list of prompts, it’s a prioritised backlog: which gaps to close first, based on commercial value and how badly the client is currently losing that prompt.

Our Prompt Research Process, Step by Step
- Pull real buyer language. Support tickets, sales call notes, review sites, and Search Console queries.
- Draft the raw prompt list. Every phrasing a buyer might realistically type into an AI engine.
- Run each prompt manually. Across ChatGPT, Perplexity, Gemini, Copilot, and Google AI Overviews. Brand missing or wrong competitor named → flagged as priority gap
- Score for intent and value. BOFU, commercial-investigation prompts outrank generic awareness prompts.
- Cluster by topic and entity. Grouped so one page can credibly answer a whole cluster, not just one prompt.
- Hand off the prioritised list. Feeds directly into citation engineering and content briefs.
What is prompt research, and why does it come before anything else?
Prompt research is the work of finding out exactly what people are asking AI engines when they’re looking for something a client sells, then checking what those engines currently answer. It’s the GEO equivalent of keyword research, but not a rename of the same task. A keyword list tells you what people type into a search box; a prompt list tells you what people say to a chat interface — and those diverge more than most teams expect. A buyer might search “best CRM for small teams” on Google but ask ChatGPT “I run a 6-person agency, what CRM should I use and why” — full sentence, context included, expecting a reasoned answer instead of ten blue links.
We run this before touching a single page because it answers the only question that matters at the start of a GEO engagement: which prompts is the client losing, and which losses are worth fixing first. Without that list, “improve AI visibility” is a slogan; with it, a backlog with a defensible order.
How does our prompt research process work, step by step?
The process runs in six steps, in the same order, every time. We skip nothing, even when a client is confident they already know their top prompts — the sources below usually surface phrasing nobody would have thought to write down.
- Pull real buyer language from the client’s own channels. Recent support tickets, sales call notes, review-site comments, and any Search Console data available. Deliberately not a brainstorm — we want words customers actually used, not words marketing assumes they used.
- Draft the raw prompt list. Every plausible phrasing gets written down — short and long, formal and conversational, comparison-style and advice-style. A single underlying question (“is X worth it”) can produce eight or ten variants, since AI engines answer near-identical questions differently depending on phrasing.
- Run each prompt manually across engines. We type every prompt into ChatGPT, Perplexity, Gemini, Copilot, and Google AI Overviews by hand and record what comes back — which sources get cited, whether the client or a competitor is named, and whether the answer is accurate. Automated tools exist, but they don’t replace reading the response text.
- Score each prompt for intent and business value. “What is GEO” is awareness-stage and low-value even if asked often. “Which AI visibility agency should I hire in India” is bottom-of-funnel and worth fixing even if asked rarely. We weight by where a prompt sits in the buying decision.
- Cluster prompts by topic and entity overlap. Prompts get grouped so one well-built page can credibly answer a whole cluster. This is also where we catch prompts no existing page can honestly answer — a content gap, not a citation-engineering task.
- Hand off the prioritised list. The client gets a ranked backlog: which clusters to fix first, why, and what page or asset each maps to.
How do we score and prioritise the prompts once we’ve found them?
Prioritisation runs on two axes: how much the prompt is worth if the client wins it, and how badly they’re currently losing it. A prompt with strong commercial intent where the client is completely absent sits at the top. A prompt with weak commercial intent where the client already gets cited occasionally sits near the bottom — less upside, less urgency.
Volume, in the search-engine sense, plays almost no role here. A prompt might get asked a handful of times a month across all AI engines combined and still be worth fixing first, because the person asking it is close to a purchase decision. That’s the practical difference between prompt research and keyword research: keyword research leans on search volume as a proxy for value, prompt research leans on funnel position because reliable volume data mostly doesn’t exist yet for AI engines.
The clients who get this wrong go and fix every prompt they can find, in whatever order the spreadsheet happened to land in. The clients who get it right fix the three prompts that sit right before a buying decision, and let the rest wait.
Palash, Founder, PalV’s DM
Where do the prompts actually come from?
Most of the useful prompts come from sources that have nothing to do with search data. Support ticket subject lines are among the richest, because they’re written in the customer’s own words under real pressure. Sales call notes are another, particularly the objections and comparison questions raised on discovery calls, since those map closely to the comparison-style prompts buyers now put to AI engines instead of a salesperson.
Review sites round this out. Comments on G2, Capterra, or category-relevant forums surface the gap between how a company describes itself and how customers describe the problem it solves — and that gap is exactly where AI engines tend to answer with a competitor’s name instead. Search Console data is useful as a starting point, but it reflects what people type into a search box, not what they say to a chat interface when they can ask a full question instead.
How is this different from ordinary keyword research?
| Keyword research | Prompt research |
|---|---|
| Built from search volume and keyword-tool exports | Built from support tickets, sales calls, reviews, and manual engine testing |
| Optimised for ranking a page in the results list | Optimised for being the source an engine cites in a generated answer |
| Short, fragment-style queries (“crm small business”) | Full-sentence, context-rich questions (“what CRM should a 6-person agency use”) |
| Success measured by position and clicks | Success measured by whether the client is named, cited, or recommended |
| Volume data is reliable and abundant | Volume data is thin or absent, so funnel position drives priority instead |
The two disciplines overlap enough that teams assume one covers the other. It doesn’t. A page built purely to rank for a keyword can still be a poor answer to the equivalent prompt, because ranking rewards relevance signals an algorithm can measure, while citation rewards being the clearest, most directly useful answer to the question asked — a stricter bar. We treat prompt research as its own discipline for this reason, running alongside our peer-reviewed GEO method rather than as an afterthought bolted onto keyword lists.
What happens to the prompt list after research is done?
The prioritised list doesn’t sit in a slide deck. It feeds two workstreams directly. Prompts where the client has relevant content but it isn’t structured for citation go to our citation engineering process, where we restructure existing pages so an AI engine can lift a clean, accurate answer without rewriting the page for humans out of the picture. Prompts where no page comes close become content briefs — new pages built to close that specific gap.
Every prompt also becomes a tracked item afterward, going into our monthly five-engine citation tracking, so we can tell the client, answer text in hand, whether the fix worked. This closes the loop with the stage that usually precedes prompt research — if you haven’t seen the AI visibility audit we run before anything else, that’s what decides whether a client needs this level of prompt-by-prompt work, or something lighter.
How often do we rerun the process?
We rerun full prompt research quarterly for active clients, sooner if something changes the underlying question set — a new product line, a pricing change, a competitor entering the market, or a client noticing new questions coming through support. AI engines also update how they answer things without warning, so a prompt fixed six months ago can quietly regress, which is why tracking matters as much as research. It isn’t a one-time audit you file away; it’s a recurring input that keeps the rest of the GEO work pointed at what buyers are actually asking right now.
Want this run on your own prompts?
If you don’t know what buyers are asking AI engines about your business, or what those engines are telling them, that’s the starting point — not content, not technical fixes.
FAQ
How many prompts does a typical prompt research round produce?
It varies by how established the client’s category is and how many products they sell. A single-service local business might end up with a few dozen prompt variants across a handful of clusters. A B2B SaaS with multiple products and personas can produce several hundred raw prompts before clustering brings it down to a manageable list.
Can prompt research be automated instead of done manually?
Parts of it can — pulling raw language from ticket exports or transcripts benefits from tooling. But running the prompts and reading what each engine answers is something we do by hand, because what we’re looking for — a wrong competitor named, an inaccurate answer, a missing citation — only shows up in the response text, not in a score.
Do you use keyword research tools at all during this process?
Occasionally, as a secondary input rather than the starting point. Search Console data or a keyword tool can suggest topics worth checking, but we always convert those into full-sentence prompts and test them directly in the AI engines rather than assuming search volume tells us anything about prompt performance.
What if a client has almost no existing content to work from?
Prompt research still runs first — it just points toward a different follow-up. Instead of citation engineering on existing pages, most of the prioritised list becomes new content briefs, since there’s nothing yet worth restructuring. The research step is unchanged; only which workstream picks up the output changes.
Does prompt research replace the AI visibility audit?
No, they answer different questions. The audit establishes the client’s overall starting position — how visible they are across engines, and roughly where the biggest problems sit. Prompt research goes a level deeper, building the specific, prioritised list of individual questions to fix, which is why it runs right after the audit rather than instead of it.
Short version: we pull real buyer language from support, sales, and reviews, draft a raw prompt list, run every prompt by hand across five AI engines, score by funnel value, cluster by topic, and hand over a prioritised backlog. That backlog decides which pages get rebuilt for citation and which gaps become new content — not a keyword export, not a guess.