Designing a 30-Prompt Test Set for AI Visibility
A practical build guide for a 30-prompt AI visibility test set: the five prompt types, how many of each, exactly how to phrase them, and the mistakes that make a set useless.


A good 30-prompt test set for AI visibility spreads across the buyer’s journey, mixes prompts you should win with prompts you don’t yet, names real competitors, and stays fixed once built so results compare month to month. This post is the practical build guide: the categories to cover, how many prompts to allocate to each, exactly how to phrase them, and the mistakes that make a test set useless. If the previous piece explained why you track, this one is the blueprint for the instrument itself.
Key takeaway
- Thirty prompts is the practical ceiling for manual tracking — enough to be representative, few enough to run several times each across multiple engines.
- Allocate across five prompt types: category, comparison, solution, use-case, and brand — so you measure the whole funnel, not one slice.
- Phrase prompts as natural buyer questions, not keywords, and lock the wording once built so months stay comparable.
Why 30 prompts, and how to think about the number
Thirty prompts is the practical upper limit for tracking AI visibility by hand, because each prompt should be run several times across several engines, and that multiplies fast. Thirty prompts, run five times each across four engines, is 600 individual checks per cycle — near the ceiling of what’s sustainable manually. Fewer than about 15 and your sample is too thin to be representative; more than 30 and you should be looking at a dedicated tool.
Think of the 30 as a portfolio, not a list. Each prompt is a probe into a different part of how buyers ask about your category, and the value comes from covering the range deliberately rather than piling up variations of the same question.
The five prompt types to include
Allocate your 30 prompts across five categories so the set measures the whole buying journey.
- Category prompts (about 6). Broad questions defining your space — “what is generative engine optimization,” “how does AI visibility work.” These test whether you’re associated with the topic at all, and they’re where educational content earns citations.
- Comparison prompts (about 8). “Best GEO agencies in India,” “top AI visibility services,” “X vs Y.” These are the highest-commercial-intent prompts and where being named directly wins business. Weight them heavily.
- Solution prompts (about 8). Specific how-to questions your service answers — “how do I get my brand cited by ChatGPT,” “how to appear in AI Overviews.” These test your practical content and often have clearer citation opportunities.
- Use-case prompts (about 5). Questions framed around a buyer’s situation — “how can a D2C brand improve AI visibility,” “AI search strategy for a SaaS startup.” These surface whether you’re recommended for specific buyer types.
- Brand prompts (about 3). Direct questions about you — “what does PalV’s DM do,” “is PalV’s DM good for GEO.” These test whether engines describe your brand accurately, which is its own visibility problem worth catching.
600 checks
What a 30-prompt set becomes when run five times across four engines in one cycle. This multiplication is exactly why 30 is the manual ceiling — beyond it, sustaining repeated runs across engines by hand breaks down and a dedicated tool becomes necessary.
Source — AI-visibility measurement practice, 2026
How to phrase each prompt
- Write them as a person would ask. AI engines are queried in natural language, so your prompts must be natural language. “Which agencies are best for getting cited in AI answers?” is a real prompt; “best AI citation agency” is a keyword that no one types into ChatGPT.
- Vary the phrasing across the set. Real buyers ask the same underlying question many ways. Include some direct (“who should I hire for GEO”), some exploratory (“how do companies improve AI visibility”), and some comparative (“what are alternatives to doing GEO in-house”). This mirrors how the question actually reaches the engine.
- Include the specifics your buyers include. Location, industry, company size, budget — if your buyers say “in India” or “for a small business,” put that in the prompts, because it changes which brands the engine surfaces.
- Name real competitors in a few comparison prompts. “Is [competitor] or [you] better for GEO” tests head-to-head positioning directly and reveals exactly how the engine frames you against a named rival.
Mistakes that make a test set useless
A few errors quietly ruin the instrument. Only including prompts you already win gives a flattering but useless read with no room to show growth. Editing prompts between cycles destroys comparability — lock the wording. Running each prompt once mistakes random variation for signal. Ignoring the engines your buyers actually use — if your audience lives in Perplexity, a ChatGPT-only set misses the point. And writing keywords instead of questions tests something the engines never actually receive.
Locking and versioning the set
Once built, freeze the 30 prompts as version one and date it. Run that exact set every cycle. When your market genuinely shifts and you need new prompts, create version two as an additive layer rather than editing version one, so your original trend line survives intact. Treat the test set like a measurement standard: its value is precisely in not changing, because only a stable instrument can show you real movement over time.
Frequently asked questions
How many prompts should an AI visibility test set have?
Around 30 is the practical ceiling for manual tracking. Each prompt should run several times across several engines, so 30 becomes hundreds of checks per cycle. Fewer than about 15 is too thin to be representative; more than 30 becomes unsustainable by hand and signals it’s time for a dedicated AI-visibility tool. Thirty balances coverage against effort.
What types of prompts should I include?
Spread them across five types: category prompts (what is X), comparison prompts (best X, X vs Y), solution prompts (how do I do X), use-case prompts (X for a specific buyer type), and brand prompts (direct questions about you). Weight comparison and solution prompts most heavily, since those carry the highest commercial intent and clearest citation opportunities.
Should I use keywords or questions as prompts?
Questions, phrased exactly as a real buyer would ask an AI engine in natural language. “Which agencies are best for AI visibility in India?” is a real prompt; “AI visibility agency India” is a keyword no one types into ChatGPT. Engines receive natural-language questions, so your test set must contain natural-language questions to measure anything real.
Can I update my test set later?
Yes, but additively. Freeze your original 30 as version one and keep running it unchanged so its trend line stays valid. When the market shifts, add a version-two layer of new prompts rather than editing version one. Treat the set like a measurement standard — its value depends on stability, so preserve the original even as you extend it.
The bottom line
A 30-prompt test set is a designed instrument, not a random list. Spread it across category, comparison, solution, use-case and brand prompts; phrase each as a real buyer question with the specifics your audience uses; include prompts you don’t yet win; and lock the wording once built. Do that and you’ve got a stable, representative gauge of your AI visibility that turns an invisible channel into a measurable one.
We design and run your custom prompt test set every month as part of our AI Visibility service.