Our AI visibility testing method is a fixed set of buying questions, run repeatedly across four engines on a stated schedule, with every result recorded as a fraction that has its denominator written down. That is the whole protocol, and this page publishes it in full so anybody can copy it, attack it, or use it to check whether our numbers deserve believing.
We are publishing the method before publishing results on purpose, and it is worth saying why, because it is not modesty.
Why the method comes first
Marketing research has no replication norm. Studies are published, quoted for years, and never re-run by anybody. Fields that do take replication seriously have found the results uncomfortable. When the Center for Open Science coordinated replications of 21 high-profile social science experiments from Science and Nature, 13 of the 21 replicated, and the replication effect sizes came in about 50 percent smaller than the originals, despite sample sizes roughly five times larger.
Those were preregistered replications of peer-reviewed work by professional researchers. The base rate for a vendor-published marketing study with no stated method, no sample size and no raw data is not going to be better than that.
So the standard we hold ourselves to is simple. If we cannot state the method, the sample and the date, we do not publish the number. Everything on this site that carries a figure will link back here.
The protocol
| Element | What we do | Why |
|---|---|---|
| Question set | A fixed list per category, written before any run, changed only with a dated note | A set that moves between runs makes the runs incomparable |
| Engines | ChatGPT, Perplexity, Claude and Google AI Mode, every time | Each maintains its own corpus, so a single engine is not a measurement |
| Runs | Minimum two per question per engine, on separate days | One check is a sample. Two on different days catches the obvious variance |
| Session state | Fresh sessions, no account history, memory and personalisation off | Otherwise you measure your own history rather than the engine |
| Recording | Full answer text stored, not a score | A score cannot be rechecked. An answer can |
| Reporting | Counts with denominators, never a composite out of 100 | A number nobody can decompose is not a measurement |
| Negative results | Published, including when they contradict what we expected | A library where everything worked is not evidence |
What we control for, and why each one matters
Sampling variance
These systems do not return the same answer twice. Part of that is the index changing underneath you and part is the model sampling probabilistically as it writes. A clinical study running identical cases through one model at five temperature settings, 10,000 outputs in total, found unique outputs rising from an average of 4.5 per case to 26.25, a 483 percent increase in divergence, with accuracy falling as the setting rose.
That was medicine rather than software procurement, and the mechanism is the same one operating when you check your brand twice and get different answers. It is the single strongest argument for repeated runs, and the reason we treat any single-run claim as an anecdote including our own.
Personalisation and memory
An account with history is not a neutral instrument. If you have searched your own company, the assistant may surface it more readily for you than for a stranger, which produces a comfortable and completely false reading. We run fresh sessions with memory disabled, and we recommend anybody replicating this does the same before concluding anything cheerful.
Question phrasing
Unaided questions and branded questions measure different things and get reported separately. Asking what the best tools are in a category tests whether you exist in the consideration set. Asking about your company by name tests whether the engine can describe you. Mixing them produces a flattering average that answers neither question.
If you want this run against your own domain rather than reading about it, our free AI visibility check uses exactly this protocol and hands back the raw answers, so you can audit the method against the output.
Known weaknesses, stated plainly
Every method has failure modes and hiding them is how a method becomes marketing. Ours has four worth knowing.
- Small run counts. Two runs per question catches gross variance and will not detect a small effect. We report counts so you can judge whether a difference is worth anything, and frequently it is not.
- Geography. Answers vary by region and we test from a limited set of locations. A result that holds for us may not hold for a buyer elsewhere.
- No access to the retrieval layer. We observe outputs, not the reasoning. When a company is absent we can often say which stage failed and we cannot prove it from the outside.
- We have a commercial interest. We sell services that this research makes look necessary. That is a real bias and publishing the raw answers is the only useful counterweight we can offer.
How to attack this
If you want to test whether our numbers hold, the fastest route is to take our published question set for a category, run it yourself with fresh sessions, and compare. Divergence is interesting and we would rather hear about it than not.
The specific things we would treat as a serious challenge are a replication producing materially different counts, a demonstration that our question set is loaded toward outcomes that favour our clients, or evidence that our session hygiene is not achieving what we claim.
We will publish corrections with the same prominence as the original. That is easy to promise before it happens, so treat it as a claim to hold us to rather than a credential.
How we choose the question set
This is the most contestable part of the method and the part most vendor research hides, so it gets the most detail here. A question set can be built to produce almost any conclusion, and a study that does not publish its set cannot be evaluated at all.
We build each set from four categories, in fixed proportions, and we write the whole thing before running anything.
| Category | Share of set | What it tests | Example shape |
|---|---|---|---|
| Unaided category | 40 percent | Whether you exist in the consideration set at all | Best tools for [job the product does] |
| Qualified category | 30 percent | Whether you appear once the buyer is specific | Best [category] for [company size or sector] |
| Comparative | 20 percent | Which set you get grouped with, and who beats you | Compare the leading [category] options |
| Branded | 10 percent | Whether the engine can describe you accurately | What is [company], and what does it cost |
The weighting is the honest part. It would be trivial to load a set with branded questions, which almost any established company wins, and report a flattering number. Most published AI visibility research does not disclose its mix, and where the mix is undisclosed the number should be assumed to be favourable to whoever paid for it.
Sets are written per category rather than per client, so the same set is used for every company in a vertical. That is what makes comparison between companies meaningful, and it also means we cannot quietly tune a set to suit a particular client.
Reading a result
A worked example makes the reporting format clearer than a description of it. Take a fictional company across a ten-question set, four engines, two runs each, which is 80 individual results.
| Outcome | Count | Share | What it means |
|---|---|---|---|
| Named in the answer | 24 of 80 | 30 percent | Present in the consideration set roughly a third of the time |
| Cited with a link | 9 of 80 | 11 percent | A passage was used as a source |
| Recommended explicitly | 3 of 80 | 4 percent | Put forward as the choice |
| Described inaccurately | 5 of 24 mentions | 21 percent of mentions | A representation problem sitting behind the visibility one |
The fourth row is the one that changes what a client does next. A company named 30 percent of the time with a fifth of those mentions carrying stale pricing does not have a visibility problem worth solving first. It has a correction problem, and adding more visibility to inaccurate information makes the situation worse rather than better.
What counts as a real difference
With 80 results, small movements mean nothing. A change from 24 mentions to 27 is well within the range that repeated runs produce on their own, and reporting it as a 12 percent improvement would be dishonest arithmetic wearing a percentage sign.
Our working rule is that we do not describe a change as a change unless it holds across both runs, appears in more than one engine, and is large enough that the direction is unambiguous. Everything else gets reported as flat, which makes for dull updates and accurate ones.
The same rule applies to our own case studies. If a client improved and we cannot separate it from variance, we say we cannot separate it from variance.
The test to apply to anybody else’s research, including ours. Ask for the question set, the run count and the raw answers. All three exist or the number is decoration. A vendor who will share the method and not the raw answers is halfway there, which is still further than most.
What we publish, and what stays private
Running this for clients produces data that is not ours to publish, so the boundary is worth stating rather than leaving to assumption.
Category-level findings get published. If we test twenty companies in a vertical and find that engines consistently hedge on security vendors, that is a finding about the category and it goes out with the method attached. Individual company results stay private unless the company agrees in writing, and we do not publish a client as a case study without before-and-after numbers they have seen and signed off.
Competitor data is the awkward case. Testing a client necessarily produces data about everybody else named in the same answers. We use it in aggregate and we do not sell it, publish it against a named competitor, or approach that competitor with it. That is a policy rather than a legal constraint, and it is the sort of thing worth asking any agency about before they start.
The tooling, since people ask
There is no proprietary platform behind this and pretending otherwise would sit badly on a page about transparency.
The question sets live in a spreadsheet. Runs are executed manually in fresh browser sessions, because automating them through an API measures the API rather than the product a buyer actually uses, and those diverge. Answers are stored as plain text with a timestamp, an engine name and a run number. Counting is arithmetic.
The manual step is the expensive part and we have not found an honest way around it. Any vendor claiming full automation of this is either using APIs and calling it consumer behaviour, or sampling far less than they imply. Both are defensible if disclosed and neither usually is.
When we re-run a test, and what forces one
A study with a date on it decays, and the honest question is how fast. Our default is to re-run any named test every six months and to publish the new figures alongside the old rather than replacing them quietly.
Four things trigger a re-run before that schedule. A vendor shipping a visible change to how answers are composed. A model release that alters behaviour we depend on. Somebody replicating our work and getting a materially different result. Or a client reporting an outcome that contradicts what we published.
The last of those is the most useful signal we get and the least comfortable, because it usually means the finding was narrower than we described. When that happens the correction goes out with the same prominence as the original, dated, with the previous figure left visible.
Anything not re-run inside twelve months gets marked as stale on the page rather than deleted. A superseded figure with a date is more useful to a reader than a missing one, and it makes it obvious when we have let something drift.
One last note on incentives, because it belongs on this page rather than buried. We publish this method partly because it is useful and partly because it constrains us. A documented protocol is harder to quietly bend when a client would prefer a better number, and a public commitment to publishing negative results is harder to abandon than a private intention to be honest. Treat that as self-interest aligned with your interest rather than as a virtue, which is a more durable arrangement anyway.
A note on why this page exists at all rather than living in a client deck. Publishing a method invites people to find fault with it, which is uncomfortable and is the entire point. A protocol nobody has attacked has not been tested, it has only been asserted, and the difference matters most in a field this young.
Frequently asked questions
Why not use an AI visibility tool instead?
Most produce a composite score you cannot decompose, and several rotate their prompt sets between reports, which quietly makes the reports incomparable. If a tool publishes its question set, its run counts and its raw answers, it is doing the same thing as this method with better ergonomics, and worth using.
How many runs are enough?
More than one, and beyond that it depends on how large an effect you need to detect. Two per question per engine is our floor for reporting anything at all. For a claim we intend to publish as a finding, considerably more.
Does it matter which account you run the test from?
Enormously, and it is the most common way a self-assessment goes wrong. An account that has previously searched your company can surface it more readily. Fresh sessions with memory and personalisation disabled are the minimum.
Why publish negative results?
Because a body of research where every test confirmed the hypothesis tells you about the publisher rather than the subject. It is also the only way our positive results mean anything.
Can I reuse this method commercially?
Yes, including if you are a competing agency. The method is not the moat and pretending otherwise would sit badly next to everything else on this page.
Studies run using this protocol are collected in the research hub as they publish, with the date and run count attached to each.
The method is only worth anything if somebody checks it.
Take the protocol, run it against your own domain, and see whether our numbers hold. Everything on this page is designed to be copied, including by people who compete with us.
If you would rather see it run once before building the habit, our free AI visibility check uses exactly this protocol and hands back the raw answers rather than a score. You keep the baseline either way.
