ChatGPT picks companies to name by rewriting your question into search queries, retrieving passages that match those queries, scoring the passages a second time, and then composing an answer from whichever ones survived. Company names surface when they appear inside passages that clear all four steps, which means you are not competing for the question a buyer typed. You are competing for a rewritten version of it that nobody shows you.
That sounds like a technicality. It is the reason most attempts to influence these answers aim at the wrong thing.
The pipeline, in the order it runs
Five steps, and a company can be eliminated at any of them for reasons that have nothing to do with the step before.
| Step | What happens | What gets you eliminated |
|---|---|---|
| 1. Reformulate | Your question becomes one or more machine-shaped search queries | Nothing yet, but the query you compete on is now different from the one asked |
| 2. Retrieve | Passages are pulled from the corpus against those queries | Not in the corpus, or no passage matches the rewritten phrasing |
| 3. Rerank | A slower, more accurate pass rescores the candidates | Retrieved but judged less relevant than the alternatives |
| 4. Compose | An answer is written using the surviving passages | Your passage does not answer anything standalone, so it contributes nothing |
| 5. Attribute | Sources are named, sometimes with links | The fact was used, and credited to a source with more corroboration |
Step three is the one people have never heard of and it explains a lot. Getting retrieved is not the same as getting used. A passage can clear the first filter and lose at the second, which is why your visibility can move without anything on your site changing.
Why you never see the query you are competing on
A buyer asks which contract tool suits a small legal team. The system does not search that string. It generates something closer to a set of separate lookups covering contract management software, features relevant to small teams, and pricing tiers, then retrieves against each.
Your page might be perfectly optimised for the buyer phrasing and match none of the rewrites. This is the mechanical reason keyword targeting transfers so poorly here, and the reason writing plainly about what something is beats writing to match a phrase. Matching happens on meaning rather than exact wording, so ambiguity costs you more than imprecise phrasing does.
The practical consequence. Test by asking the question a buyer would ask, then test again using the flattest, most literal phrasing of the same need. If you appear for one and not the other, you have a retrieval problem rather than a presence problem, and the fix is in how plainly your passages state things.
What actually decides which names appear
Four factors dominate, and they are not equally weighted.
- Whether you are in the corpus at all. Binary, upstream of everything, and the most common failure. Nothing downstream matters if this fails.
- Whether a passage answers the rewritten query standalone. This is the part you control most directly through editing.
- Whether your name resolves to a distinct company. An ambiguous name makes naming you risky, and hedging costs the system nothing.
- Whether anybody else says the same thing. A claim corroborated off your domain is safe to repeat. A claim that exists only on your site is not.
The fourth is where most of the effect sits and the least of the effort goes, because it is the only one that cannot be done from inside your own content management system.
Why the same question names different companies on different days
Two causes stack, and neither is a sign that something changed on your end.
The corpus moves. Pages get recrawled, new material appears, and the candidate set for a query is not fixed. And generation samples probabilistically, so even an identical candidate set can produce a different composition.
The practical consequence is that a single check is a sample rather than a measurement. If you look once, see a competitor, and conclude you have a problem, you may be reacting to noise. If you look once, see yourself, and conclude you are fine, the same applies with worse consequences.
Distinguishing a real change from ordinary variance needs repeated runs across engines, which is tedious enough that almost nobody does it properly. Our free AI visibility check runs it and hands back the raw answers rather than a score.
What you can influence, honestly ranked
| Lever | How much control you have | Speed |
|---|---|---|
| Crawler access and rendering | Total. It is a configuration setting | Days |
| Passage extractability | High. It is editing your own pages | Weeks |
| Entity resolution | High, and it depends partly on third-party records | Weeks to months |
| Third-party corroboration | Low. Other people decide | Quarters |
| Reranking behaviour | None. It is a product decision by the vendor | Not applicable |
Three things people get wrong about this
That there is a ranking to climb. There is no position. There is a set of passages that survived a filter and a composition step that used some of them. Optimising for a rank that does not exist produces reports about a number nobody can act on.
That the model is what you influence. The model composes. The retrieval system in front of it decides what the model sees. Almost everything you can influence sits on the retrieval side, which is why the unit of competition changed from a page to a passage.
That more content helps. Volume adds candidates for retrieval and does nothing for reranking or attribution. Forty mediocre pages lose to one page containing a checkable number in the same sentence as its claim.
How to work out where you are being eliminated
- Ask about your company by name in a fresh session. Blank or wrong means you are failing before retrieval and the problem is identity.
- Ask a category question and see whether competitors appear fluently. If they do and you do not, the corpus covers your space and you specifically are missing from it.
- Ask a question your best page directly answers, using flat literal phrasing. Appearing here but not for the natural phrasing points at retrieval rather than presence.
- Search a statistic or framing you originated. If it comes back credited to somebody else, you are being used at step four and losing at step five.
Each of those isolates a different step, and all four take about twenty minutes together. Which one fails first determines whether this is an engineering ticket, an editing job or a slow reputational project.
When it answers from memory instead of looking anything up
Not every answer triggers a search. For a question the model considers well established it may answer from what it absorbed during training, and that path has completely different properties from the retrieval path described above.
Training data is a snapshot. It has a cutoff, it cannot be corrected by publishing something today, and it may describe a version of your company that stopped existing two funding rounds ago. If an assistant states your pricing confidently and gets it wrong with no sources shown, this is usually why.
You can often tell which path produced an answer by whether sources are cited at all. Cited sources mean retrieval ran and you can influence the outcome. No sources and a confident tone means the answer likely came from memory, and the only remedy is publishing enough current, well-corroborated material that future retrieval overrides it.
This is also why blocking training crawlers costs you nothing in today’s answers but does affect what a future model absorbs. It is a decision about the next model generation rather than this quarter, which is a reasonable trade either way as long as it is made knowingly.
Worth adding one caution about all of this. The pipeline described here is how these systems broadly work, and the specifics are product decisions that change without announcement. Reranking behaviour, how aggressively queries are rewritten and when retrieval fires at all have all shifted more than once. Build on the parts that are structural, which are that retrieval works on passages, that corroboration lowers risk and that ambiguity is expensive. Those hold across implementations. The rest is weather.
What this means for how you write
Everything above collapses into a short editorial instruction, which is that the sentence is the unit rather than the page.
A passage gets retrieved because it matches the meaning of a rewritten query, survives reranking because it is clearly about that thing, and gets used because it answers something on its own. All three properties are decided inside individual sentences, not by the structure of the document containing them.
Which is why the most effective edit is usually deletion. Removing the sentence that says as we saw above, replacing this and that with the actual noun, and moving the number into the same sentence as the claim it supports does more than adding two thousand words of new material.
The uncomfortable summary is that the two steps that matter most sit at opposite ends of the difficulty range. Crawler access is trivial and skipped. Corroboration is hard and decisive.
Related reading
- Why AI Overviews cite some brands. The same question for Google surfaces, where the answer differs in useful ways.
- Indexed, retrieved, cited. The three stages these five steps collapse into, and how to tell them apart.
- Mention, citation, recommendation. Being named, sourced and chosen are three outcomes with three different causes.
