The AI Visibility Research Hub: Our Methods and Every Study We Publish

The standard every study here has to meet, published before the first result rather than after it. Fixed question sets, stated run counts, named engines, and the findings that go against us.

The AI visibility research published here covers original tests of how answer engines choose which companies to name, and this page is the index of every one of them along with the standard each has to meet before it goes out. Right now that index is short, because the standard was written before the first result, which is the entire point of publishing this page at slot nine rather than at slot ninety.

Almost all research in this field is produced by companies selling services against its conclusions, and we are one of them. There is no ethics board and no peer review. So the only thing separating a study from a sales asset is whether the method was fixed in advance and whether the numbers can be recounted by somebody who disagrees with you.

What follows is our publishing standard, stated first so you can hold us to it.

The replication problem, and why this field is worse

Even in disciplines with peer review, formal methods training and career penalties for getting it wrong, findings do not hold up as often as people assume. A team led by the Center for Open Science re-ran 21 high-profile social science experiments originally published in Science and Nature between 2010 and 2015, using samples roughly five times larger than the originals and preregistering every protocol. Thirteen of the 21 replicated, and the effects that did survive were about half the size originally reported.

Sit with that for a second. Those are papers that cleared the highest bar publishing offers, and a third of them did not hold. Marketing research has none of that infrastructure. No preregistration, no replication norm, no independent review, and a direct commercial interest in the result.

Which is why we assume our own findings will shrink under scrutiny and design for it, rather than presenting each one as settled.

If you only want your own numbers

This page is about how we run studies. If what you actually need is a measured baseline for your own domain rather than category research, a free AI visibility check produces one, and you keep the raw answers whether or not you read another word of this.

The five things every study here states

This is the standard. If something we publish is missing any of these five, it is a mistake and worth telling us about.

#The standardWhy it is thereHow to check we did it
1The question set is published in full, and was fixed before collectionA prompt set adjusted mid-study can produce almost any conclusionEvery study links its full question list
2Run counts and dates are stated, never a single passAnswers vary between runs, so one pass is a sample rather than a measurementLook for the denominator on every figure
3Every engine tested is named, including the ones that went nowhereReporting only the engine that produced a clean result is selection by another nameEngine list appears in the method section
4Findings that contradict our commercial interest are published anywayAn agency publishing only wins is publishing advertisingLook for the studies where nothing moved
5Raw answer sets are available on requestA number nobody can recount is an assertionAsk us. We will send them
Written before the first study, so it cannot be quietly relaxed to fit a result we liked.

What we will not publish, and why

Three formats are common in this space and none of them survive the standard above.

A composite visibility score. A single number out of a hundred, produced by a formula nobody outside the vendor can inspect. It moves, which feels like information, and it cannot be decomposed into anything you could act on. We report counts with denominators instead, because those can be argued with.

Single-run comparisons. Screenshots of one answer, presented as evidence of a trend. Given how much output varies between identical runs, a single pass tells you what happened once and almost nothing about what happens generally.

Undisclosed prompt sets. If the questions are secret, the result cannot be checked, and the temptation to pick questions that flatter the conclusion is enormous. Every set we use is published, including the questions where we performed badly.

The four kinds of study, and what each one can prove

Not every study answers the same shape of question, and the commonest way research in this field misleads is by presenting one kind as though it were another. A single product comparison is not evidence about a market. A correlation is not a cause. Stating which kind you are reading is part of the standard.

KindWhat it doesWhat it can establishWhat it cannot
Named product testRuns a fixed question set about one category and records which vendors are namedWho is currently visible for those questions, on those engines, on those datesWhy, or whether it will hold next quarter
Correlation studyCompares a measurable property against citation frequency across many companiesWhether a property moves with visibilityThat the property caused it
Index editionRepeats an identical measurement on a scheduleDirection of travel over time, which is the only way to see a trendAnything about companies outside the sample
Vertical benchmarkMeasures one category deeply, including engine-by-engine differencesWhat normal looks like in that category, which is what most companies lackThat another vertical behaves the same way
Every study published here states which of the four it is in the first paragraph.

The distinction that gets abused most is the second row. A correlation study finding that companies with more reviews are cited more often does not establish that getting reviews causes citations, because larger and better-known companies tend to have both. Saying so in the study rather than in a footnote is the difference between research and a brochure.

How the question sets are built

A study is only as good as its questions, and question selection is where bias enters most easily and least visibly. Four rules govern how ours are built.

  • Questions describe a job, not a category. Buyers who do not know a market yet describe a problem. Testing category names measures vendors who already won the naming argument.
  • The set includes questions we expect to lose. A set assembled from questions where our clients perform well produces a flattering and worthless result.
  • No brand names in unaided questions. Naming a company guarantees it appears, which measures the model’s agreeableness rather than anything about visibility.
  • Frozen before collection starts. Published in full with the study, so anybody can rerun it and disagree with what we found.

That last rule is the one with teeth. Once the set is published and the collection has begun, adding a question because the early results looked disappointing is not an option, and the temptation to do exactly that is the reason the rule exists.

How to read anybody’s AI visibility research, including ours

This part is worth having whether or not you ever read a study of ours, because the same four questions dismantle most of what circulates in this field.

  1. What is the denominator? Forty percent of what? If the answer is not a count of runs across a stated question set, the percentage is decoration.
  2. How many times was each question run? Once is an anecdote. The number should be stated without you having to ask.
  3. Which engines, and were any excluded after the fact? A study covering only the engine where the finding was strongest is not a study of the market.
  4. Who benefits if this is true? Applies to us as much as anyone. It does not make a finding wrong, and it tells you how hard to look for the caveats.

Apply those four to the next AI visibility statistic you see quoted in a pitch. Most of them do not survive the first one.

A fifth question is worth adding when the claim is about change over time. Was the measurement identical on both occasions, or did the question set, the engines or the session conditions shift between them? A great deal of reported improvement in this field is an artefact of measuring something slightly different the second time, and it is rarely deliberate. It is just that nobody wrote the method down precisely enough to repeat it, which is the failure mode this entire page exists to avoid.

How many runs count as enough

The question we get asked most, and the honest answer is that it depends on how large a difference you are trying to detect. A finding that one vendor appears in nine of ten runs and another in one of ten needs far less sampling than a claim that one appears slightly more often than another.

The rule we work to is that a single run is never reportable, that a claimed difference smaller than the run-to-run variation we observed is not a difference, and that a study covering fewer runs than it has conclusions is overreaching. Those are unglamorous constraints and they eliminate most of the confident statistics circulating about this field.

For a model of what disclosure should look like, the Pew Research Center’s work on AI summaries and click behaviour is worth reading purely for how it is written up. It states the panel, the month, the exact number of searches analysed and the subset that produced summaries, before it states a single finding. You can disagree with the interpretation while trusting the numbers, which is the whole objective.

Marketing research almost never does that, and the omission is rarely accidental. A figure without a denominator cannot be argued with, and being un-arguable is commercially useful in a way being correct is not.

What is running now

Being concrete about the state of things, since a research hub with no studies on it should say so plainly rather than implying otherwise.

Continuous collection is running across the major answer engines against fixed question sets in several B2B software categories. The first index edition publishes once a full quarter of collection sits behind it. Named product comparisons and the correlation studies follow after that, because both need a longer baseline than a single quarter provides.

Every one of them will appear on this page as it goes out, with the method section attached and the raw answers available. If the first index shows that visibility barely moved for anyone, that is what it will say.

There is a commercial cost to working this way and it is worth naming, because it explains why most agencies do not. A study that concludes an intervention did nothing is difficult to sell against. Publishing the question set means competitors can rerun it. Handing over raw answers means a client can check our arithmetic. Each of those is a small act of disarmament, and collectively they are the only reason to believe anything on this page.

What a published study will contain

So there is no ambiguity about the format, every study here carries the same six sections in the same order.

  1. The finding, stated as a count with its denominator, in the first sentence rather than after a preamble.
  2. What kind of study it is, drawn from the four above, with what that kind cannot establish.
  3. The full question set, published rather than summarised.
  4. Method, covering engines, dates, run counts, session conditions and anything excluded.
  5. What would change our mind, written before we saw the result.
  6. Limitations, in the body rather than in small print at the end.

The fifth one is unusual and it is the section we find hardest to write. Committing in advance to what evidence would overturn a finding removes the option of explaining away an inconvenient result later, which is precisely why it belongs in there.

Frequently asked questions

When will the first study be published?

Collection runs continuously and the first index edition publishes once it has a full quarter behind it. We would rather be late than publish a number built on three weeks of thin sampling, because the whole point of this page is that we said what the standard was before we had a result to defend.

Why publish the method before any results?

Because a standard written after you have seen your data is not a standard, it is a rationalisation. Stating the question set, the run count and the engines in advance removes the option of quietly adjusting any of them later to make a finding look better.

Will you publish studies that make AEO look ineffective?

Yes, and we expect to. A test where the intervention did nothing is more useful to a buyer than a test where it worked, because it narrows what is worth paying for. An agency that only ever publishes wins is publishing marketing.

Can I reuse your data?

Yes, with attribution. Raw answer sets are available for anything we publish. If somebody recounts our numbers and reaches a different conclusion, that is the system working rather than a problem.

Do you accept sponsorship or vendor participation in studies?

No, and there is no version of it we would accept. A vendor paying for inclusion in a comparison has an interest in the result, and no disclosure statement repairs that. If a company wants to know how it performs against a question set, that is an audit rather than a study, and we keep the two entirely separate.

What happens when a study contradicts something you published earlier?

We publish the contradiction and leave the original up with a dated note pointing at it. Quietly editing a previous finding to agree with a newer one destroys the record that makes any of this checkable, and this field moves fast enough that being wrong six months later is the expected case rather than an embarrassment.

How is this different from the visibility reports vendors publish?

Mostly in what gets disclosed. The common pattern is a composite score with no denominators, no run count and an undisclosed prompt set, which cannot be checked or reproduced by anyone. Every number here is a count with the denominator attached.


For the mechanics these studies measure, the complete guide to AEO covers the five gates between content and a citation.

The standard above is also a reasonable one to hold your own reporting to.

Fixed question set, stated run counts, every engine named, denominators on every number. If your current AI visibility reporting cannot produce those four things, the reporting is the thing to fix before the visibility is.

A free AI visibility check gives you a baseline built to this standard, with the raw answers handed over as yours to keep. The newsletter is where each study goes out first.

Shaban Asif, founder of Uncited Brands

Founder, Uncited Brands

Shaban Asif

Shaban runs answer engine optimisation for B2B SaaS companies at Uncited Brands. He works the unglamorous end of the problem, mostly crawler access, retrieval diagnostics and measurement baselines, and publishes the tests that do not go his way alongside the ones that do.

Connect on LinkedIn