A prompt set is a fixed list of questions you run against AI engines on a schedule, logged out, recording every answer word for word. Twenty questions covers it. The set never changes, which is the entire point, because answers only mean something when they are comparable to the last time you asked.
Run this before you hire an agency, including us. It costs an afternoon and it changes what you are buying. Roughly a third of the companies that run it properly discover the problem is something their existing team can fix in a fortnight, and that is a cheaper outcome than any retainer.
The four rules that make the results mean anything
Most people run a few prompts, get an answer they dislike, and conclude something. The conclusion is usually wrong, because they varied five things at once and measured none of them.
- Logged out, private window, every single time. A logged in session carries memory, chat history and personalisation. It will flatter you. Your buyers are not logged into your account.
- Four engines, not one. ChatGPT, Google AI Mode or AI Overviews, Perplexity and Claude. They read different indexes and disagree constantly, and the disagreement is data rather than noise.
- Two runs per prompt per engine, at least an hour apart. These systems are not deterministic. A single run tells you what happened once. Two runs tell you whether it was stable.
- Paste the answer verbatim into a sheet. Not a summary, not a score, not your impression of it. The exact text, with the date and the engine. Your memory of what it said last month is worthless and you will trust it anyway if you do not have the transcript.
That is one hundred and sixty answers if you do the whole thing, which sounds enormous and takes about three hours because most of it is pasting. Write the protocol down before you start rather than after, which is a habit borrowed from preregistration in research and exists for one reason. It stops you quietly adjusting what counts as success once you have seen the results.
The discipline that matters most. Decide what a pass looks like for each prompt before you run it. Write it in the sheet. Otherwise you will read an answer that mentions you in passing at position nine and record it as a win, because you wanted a win.
Group A: does it know you exist
Four prompts establishing whether there is anything to work with. If this group fails, nothing in the other sixteen matters yet, and any agency proposing content strategy before fixing this is selling you the wrong thing.
1. What is [your company]?
2. Who founded [your company] and what do they do?
3. Is [your company] a legitimate company? What do you know about them?
4. What does [your company] sell, and who is it for?
You are looking for three failures here. Nothing at all means you are blocked, rendering client side, or too new to have been crawled. The wrong company means you have an entity resolution problem and are being confused with a similarly named business. Vague but not wrong, the answer that could describe any company in your category, means you have been indexed but nothing distinctive about you has been retained.
That third one is the most common and the most expensive to ignore, because it looks like success in a screenshot. We went through the full diagnostic tree for this group in why your SaaS is not showing up in ChatGPT, which is worth having open while you read the answers.
Group B: does it name you unprompted
This is the group that matters commercially, because it is the only one that simulates a buyer who has never heard of you. Do not name yourself in these prompts. The moment you do, you have changed the question from do they recommend you to can they describe you.
5. What are the best tools for [the job your product does]?
6. I need [specific outcome] for a [company size] team. What should I look at?
7. Who are the leading companies in [your category]?
8. What [category] software should I shortlist if [your strongest qualifier]?
Prompt eight is the one people get wrong. The qualifier should be the thing you are genuinely best at and narrow enough to exclude most of the market. If you sell compliance automation for medical device firms, the qualifier is medical devices, not enterprise. Broad qualifiers return the same five incumbents to everybody and teach you nothing about yourself.
When Group B returns the same four names every time
Four competitors appearing across all four engines with no sign of you is not a general failure. It is one of three specific causes, and they have different fixes and very different costs.
We run this group at a larger sample as part of a free visibility check, and hand back the transcripts so you can check the reasoning rather than take the conclusion on trust.
Group C: how you fare head to head
Four prompts where you name a competitor. These surface the comparison narrative circulating about you, which is often built entirely from third party pages you have never read.
9. [Your company] vs [closest competitor]. Which should I choose?
10. What are the main alternatives to [largest competitor in category]?
11. Why do companies switch away from [your company]?
12. What do people dislike about [your company]?
Eleven and twelve are uncomfortable and they are the most useful prompts in the entire set. Whatever comes back is the objection your sales team is already fighting, now stated by a neutral sounding source at the exact moment a buyer is forming an opinion. If the answer cites a four year old review site thread, you have found something worth a quarter of work.
Prompt ten checks whether you appear in your competitor’s alternatives list, which is frequently the cheapest visibility available in the whole category and is decided almost entirely by third party pages.
Group D: the buying stage
Four prompts covering what an engine says once somebody is close to a decision. This is where pricing, integration and implementation questions live, and where thin documentation gets punished.
13. How much does [your company] cost?
14. Does [your company] integrate with [your most requested integration]?
15. How long does it take to implement [your company]?
16. Is [your company] SOC 2 compliant, and where is data stored?
Expect these to fail more than the others. The information usually exists on your site and sits in a format that survives chunking badly, most often a comparison table, an accordion that loads on click, or a PDF. When an engine says it does not know your pricing and your pricing page has been public for three years, you have learned something specific about how your page is being read rather than something vague about authority.
| What comes back | What it usually means | Where to look |
|---|---|---|
| No information at all | The page is not being retrieved or not being rendered | Server logs for the crawler, then the page in a text only browser |
| Outdated but confident | An old page or third party listing is outranking your current one | Search the old figure, find what still publishes it |
| Correct but hedged heavily | Retrieved, but the passage is not decisive enough to state plainly | The exact paragraph. It probably needs to say the number in a sentence |
| Correct and plainly stated | Working as intended | Nothing. Record it as the baseline and move on |
Group E: is it current
Four prompts checking freshness. This group catches the failure that damages deals fastest, which is an engine confidently describing a version of your company that stopped existing eighteen months ago.
17. What is new from [your company] recently?
18. Has [your company] raised funding? How much and when?
19. Who is the CEO of [your company]?
20. What industries does [your company] serve?
Nineteen looks trivial and is not. Leadership changes propagate through these systems slowly and unevenly, and an engine naming your previous chief executive to a prospect is a small credibility event that nobody on your team will ever hear about. Twenty catches positioning drift, which is when the market you served three years ago is still the market the model thinks you serve.
The result almost nobody expects
The most common finding across companies that run this properly is not absence. It is accuracy with staleness, an engine describing you fluently and slightly wrongly. It reads as a pass in a screenshot and functions as a loss in a deal, because a prospect has no way to know the detail is two years old.
Our free AI visibility check records the exact wording across engines and dates it, so you can see which version of your company is actually circulating rather than which one you published.
Scoring it without fooling yourself
Resist building an index. A single number out of a hundred feels satisfying and destroys the information you collected, because the four things being averaged have different causes and different fixes.
Score each answer on three binary questions instead. Were you named. Was what it said accurate. Was it attributed to a page you control. Three columns, yes or no, no partial credit. Then count.
| Pattern across the 20 | Reading | First move |
|---|---|---|
| Named rarely, accurate when named | Retrieval problem, not a reputation problem | Passage structure and crawler access |
| Named often, frequently inaccurate | Reputation and freshness problem | Find and correct the source being cited |
| Named in Group A only | You are a known entity with no category association | Category pages and third party corroboration |
| Named in Group B but not Group A | Something is odd. Usually a name collision | Entity resolution and disambiguation |
| Strong on one engine, absent on others | Index coverage, not content quality | Check crawler access per engine, individually |
The distinction between being retrieved and being credited runs underneath all of this, and it is the single most misread result in the set. The mechanics are set out in indexed, retrieved and cited, and the difference between the three decides which of the five rows above you are actually in.
Why run it before hiring rather than after
Three reasons, and the third is the one that saves money.
You get an honest baseline that predates anyone’s incentive to show improvement. A baseline produced by the agency you are paying, in month one, is a number with an interested party attached to it. Yours is not.
You learn the vocabulary. Walking into a pitch having read forty verbatim answers about your own company changes the conversation completely. You will ask about retrieval and attribution instead of asking about rankings, and you will notice immediately when somebody answers a different question than the one you asked.
And you may find you do not need to hire anybody. If Group A is clean, Group B fails, and every Group D answer is hedged, your problem is probably four pages that need rewriting, which is a job your existing content person can do. The fuller version of that decision, including the cases where an agency genuinely is the answer, is in the complete guide to choosing an AEO agency.
Four ways people ruin this
Running it once. One run is an anecdote. These systems vary between identical queries, and without a second run you cannot tell a real change from ordinary variance. Anyone who has looked at how sample size relates to detecting a difference will recognise the problem. Two runs is a compromise, not a rigorous protocol, and it is enormously better than one.
Changing the prompts between rounds. The moment you improve the wording, you have lost comparability with every previous round. Keep a second, separate sheet for experimental prompts. The twenty never change.
Letting the person who owns the outcome do the scoring. Scoring is where optimism enters. Formal research handles this with blinding and with published reporting standards. You do not need any of that machinery. You need the scoring done by somebody who does not report to whoever is accountable for the number.
Testing only ChatGPT. It is the one everybody reaches for and it is a single index with a single set of habits. A company invisible in ChatGPT and prominent in Perplexity has a very different problem from one that is invisible in both, and you cannot tell those apart without running all four.
How often, and what to expect
Quarterly is right for most companies. Monthly if you are actively working on this and want to see movement, which you mostly will not for the first two months. The pipeline from publishing to being retrieved to being treated as a consensus source runs on a timescale of weeks, and the last part is the slowest because it depends on other people.
The protocol above borrows its structure from systematic review, where the value comes from fixing the method before you look at the results rather than from any individual observation. Twenty prompts run carelessly tell you nothing. The same twenty run the same way four quarters running will tell you exactly when something changed and, usually, what you did the month before it did.
Related reading
- The five minute, one prompt version. If three hours is not happening this week, start here instead.
- A free way to check if AI has heard of you. The Group A question on its own, with the full diagnostic tree.
- How retrieval actually picks sources. The mechanism behind almost every failure pattern in the scoring table.
- How ChatGPT picks which companies to name. Specifically why Group B behaves differently from Group A.
- What AEO is. The wider picture, if this is your first piece of ours.
Run it yourself first. Genuinely.
Everything above is free, takes an afternoon, and produces a baseline nobody can massage later. You will walk into every agency conversation after it knowing more than most of the people pitching you, which is a reasonable return on one afternoon.
When the results are clear but the cause is not, send us the domain and we will run a visibility check. That is the same set at a larger sample, plus the server log read that tells you whether crawlers are even reaching you, and name one binding constraint. You keep all of it regardless of what happens next.
