Indexed, Retrieved, Cited: Three Stages Most Teams Collapse Into One

A company can be indexed everywhere, retrieved often and still never cited. Three stages, three unrelated causes, and the fixes have almost nothing in common.

Indexed, retrieved and cited are three separate stages, and a company can pass the first two and fail the third indefinitely. Indexed means a system holds a copy of your page. Retrieved means it pulled that page into the working set for a specific question. Cited means it used your passage in the answer and said so.

Most teams treat these as one thing called visibility, which is why so much AEO spending lands in the wrong place. The three stages fail for unrelated reasons and the fixes have almost nothing in common.

The three stages, properly separated

StageWhat it meansWhy it failsWhat fixes it
IndexedA copy of the page exists in the corpus the engine queriesCrawler blocked, page renders empty, or the URL was never discoveredAccess, rendering, sitemaps. Engineering work
RetrievedThe page enters the candidate set for one specific questionThe passage does not match the reformulated query, or the entity is ambiguousPassage-level clarity and entity resolution
CitedA passage is used in the answer and attributed to youNothing standalone to quote, or no corroboration for the claimEditing, and independent coverage
Three stages, three unrelated causes. The symptom for all of them is the same absence.

The order matters because it is strictly sequential. You cannot be retrieved without being indexed and you cannot be cited without being retrieved. Work spent on a later stage while an earlier one is broken produces exactly nothing, and it produces nothing silently, which is what makes it expensive.

The collapse that costs the most. A team sees no citations, concludes the content is weak, and commissions a rewrite. If the real failure is at indexing, the rewrite changes nothing and the conclusion drawn afterwards is that AEO does not work. The diagnosis was wrong, not the discipline.

Why being indexed tells you so little

Indexing is necessary and close to worthless on its own. Every engine maintains its own corpus and they overlap unevenly, so being present in one says almost nothing about the others. A large-scale comparison of the sources retrieved by traditional and generative search found that the two diverge significantly in source selection, domain typology and information freshness.

The practical read is that ranking well in Google is weak evidence about ChatGPT, and being absent in one engine while present in another is a corpus fact rather than a content fact.

Retrieval is where most of the confusion lives

Retrieval operates on passages and on a rewritten version of the question, neither of which you can see. Your page is split into fragments, the fragments are compared against a query nobody showed you, and the ones that match get pulled forward.

Two consequences follow. A page that ranks nowhere can be retrieved constantly because one paragraph is unusually clean. And your best page in ordinary search can be invisible here because it is written as a continuous argument, which is exactly the property that makes a fragment useless on its own.

Matching happens on meaning rather than exact wording, which is why hitting a keyword precisely matters far less than being unambiguous about what you are describing. Vector search is the mechanism, and most production systems combine it with keyword matching rather than replacing one with the other.

Working out which of the three stages is failing is the entire diagnosis, and it does not need a tool. If you would rather have it done properly against your own domain, that is what our free visibility check reports.

Citation is the only one that is partly social

The first two stages are technical and yours to fix. The third is not entirely. A passage gets cited when it answers something standalone and when the claim is safe to repeat, and safety usually comes from the claim existing somewhere other than your own site.

That is why a company can be indexed everywhere, retrieved often, and still never named. The machinery is working. The evidence is thin.

How to tell which stage you are failing

  1. Indexed? Count AI crawler hits in your server logs over thirty days. Zero means stage one. This is the cheapest check available and almost nobody runs it.
  2. Retrieved? Ask an engine a question your page directly answers, using the phrasing a buyer would use. If competitors appear and you do not, the page exists and is not being selected.
  3. Cited? Ask the same question several times across days. Appearing sometimes and not others is a retrieval and corroboration problem rather than a technical one.

Whichever check fails first is where the budget belongs, and the earlier the failure, the cheaper the fix. That relationship holds often enough to plan around.

The two handoffs where things quietly break

The stages themselves are reasonably well understood. The failures cluster at the joins between them, and both joins fail without producing any signal you would notice.

Between indexed and retrieved sits the question of whether your page is a plausible answer to anything specific. A page can be perfectly indexed and never surface, because it is about a topic rather than about a question. Nothing is broken. It simply never becomes a candidate.

Between retrieved and cited sits the question of whether a passage can stand on its own. A page can be pulled into the working set repeatedly and contribute nothing, because every paragraph depends on the one before it. The system read you and had nothing liftable.

What you observeThe tempting diagnosisWhat it usually is
No citations anywhereThe content is not good enoughNot indexed, or blocked at the crawler
Cited by one engine, absent from anotherInconsistent qualityDifferent corpora, which is a coverage fact
Named for your brand, absent for category questionsWeak authorityRetrieval, since nothing matches the reformulated query
Quoted often, recommended neverInsufficient content volumeNo corroboration off your own domain
Results swing between checksSomething changedYou are reading one sample rather than a measurement
Every row on the left is compatible with several causes. The middle column is what teams usually conclude.

Why the sequence decides the budget

Each stage costs roughly an order of magnitude more to fix than the one before it, and the effect arrives more slowly.

Indexing problems are configuration. A robots.txt line, a rendering fix, a firewall rule. Hours of work, and the change shows up within days because the crawler was already trying to reach you.

Retrieval problems are editorial. Rewriting the paragraphs that should win so they answer something standalone. Weeks of work, and the effect appears over a month or two as pages are re-crawled and re-chunked.

Citation problems are reputational. Getting other people to publish claims that corroborate yours. Quarters of work, mostly outside your control, and no guarantee at the end of it.

Which is why diagnosing the stage before spending is not process for its own sake. Getting it wrong by one stage means paying ten times more than necessary and waiting ten times longer to find out it did not work.

A useful rule of thumb. If somebody proposes a fix and cannot tell you which of the three stages it addresses, they have not diagnosed anything. That question is fair to ask in a first call and it separates a method from a package.

What to do this week

The sequence is short enough to run in an afternoon, and the first step resolves the question for a surprising share of companies.

  1. Count AI crawler hits in your access logs for the last thirty days, broken down by agent and status code. Zero hits, or a wall of 403s, ends the investigation at stage one.
  2. Pick the page you believe should win and read its strongest paragraph with nothing above or below it. If it needs its surroundings to make sense, you have a stage two problem and you now know which paragraphs to edit.
  3. Ask four engines the same category question twice, several days apart, and write down who gets named. Consistent absence with healthy logs points at stage three.

Whichever step fails first is your answer, and you have not spent anything to get it. That is the whole argument for diagnosing before buying.

Why this is worth naming explicitly

The reason to keep these three words separate is that they map onto three different budgets and three different owners. Indexing belongs to engineering and costs hours. Retrieval belongs to whoever edits the pages and costs weeks. Citation belongs to marketing and product and costs quarters.

Collapsing them into visibility hands the whole problem to one team, usually the one that owns content, and asks them to solve two thirds of it with tools they do not have. That is how a reasonable programme produces nothing and everybody concludes the discipline is fraudulent.

When somebody reports that AI visibility is flat, the useful next question is not what we should publish. It is which of the three stages the flatness is happening at, because the answer determines who should even be in the room.


One more thing worth saying plainly. If your name does not resolve to a distinct company, stages two and three both degrade at once, which makes entity resolution look like a content problem when it is not.

Related reading

Shaban Asif, founder of Uncited Brands

Founder, Uncited Brands

Shaban Asif

Shaban runs answer engine optimisation for B2B SaaS companies at Uncited Brands. He works the unglamorous end of the problem, mostly crawler access, retrieval diagnostics and measurement baselines, and publishes the tests that do not go his way alongside the ones that do.

Connect on LinkedIn