Choosing an AEO agency comes down to whether they can tell you which specific thing is broken before they tell you what they would sell you, because the work splits into technical access, entity resolution, extractability and third-party corroboration, and those need different skills and cost different money. An agency that quotes a single monthly figure without asking to see your server logs is pricing the category rather than your problem.
Almost everything written about hiring for this is written by people selling it, and we are aware that includes us. So this guide is built to be useful if you never hire anyone, including the parts that are inconvenient for our own pitch.
Start by knowing what is actually achievable
Before you evaluate anybody, calibrate what success can look like, because the honest ceiling is lower than most proposals imply and the reason is structural rather than a matter of effort.
Research published in June 2026 by Xi Chu and Yupeng Hou tested brand recommendation behaviour across GPT-4o-mini, Claude Sonnet and Gemini 3 Flash. When products had identical specifications, the established brand was recommended 100 percent of the time. Incumbency is not a tiebreaker in these systems. It is close to the whole decision when nothing else separates the options.
The same paper contains the more useful finding. That dominance disappeared once a competitor had a rating advantage of even a tenth of a star. The lock is real and it is brittle, and what breaks it is independent evidence that you are better rather than anything you say about yourself.
Two caveats. The study used skincare, where buyers cannot assess quality before purchase, with robustness checks on other product types, so it is not a B2B software finding. And it measures recommendation rather than mention or citation. The mechanism is what transfers, and the mechanism says your review profile is load-bearing in a way most software marketers do not treat it.
What to take from that into a sales call. If an agency promises to make you the recommended option in a category with an entrenched leader, and their plan is content, ask how content overcomes a hundred percent incumbency preference. The honest answer involves third-party evidence and takes quarters. Anyone who does not raise that has not read the research.
Twelve questions, and what the answers tell you
Take this list into the call. The pattern that matters is not whether they have answers, it is whether the answers get more specific under pressure or less.
| # | Ask them | A good answer sounds like | A bad answer sounds like |
|---|---|---|---|
| 1 | Which of my problems is technical and which is editorial? | We cannot know until we read your logs and robots.txt | Both, we do a full programme |
| 2 | What will you look at in the first week? | Crawler access, rendering, entity resolution, a baseline | A content audit and keyword mapping |
| 3 | Which AI agents will you check individually? | Names nine agents across four companies unprompted | We optimise for AI search generally |
| 4 | What does Google’s own documentation say about AEO? | That its SEO best practices already cover its AI features | Google has a separate AI algorithm now |
| 5 | What schema do I need for AI Overviews? | None specific, Google states this explicitly | We implement proprietary AEO schema |
| 6 | How will you measure this? | A frozen prompt set, repeated runs, counts with denominators | A visibility score out of a hundred |
| 7 | How many runs before you report a change? | Enough to separate a real move from sampling variance | We report monthly from the dashboard |
| 8 | What would make you tell me not to hire you? | A specific scenario they have actually walked away from | We can help any business |
| 9 | Show me work where the result was negative or flat | A real example with the reason it did not move | Only wins in the case study library |
| 10 | Who does the work, and are they on this call? | Named practitioners with defined hours | Our team of specialists |
| 11 | What happens to the measurement setup if I leave? | You keep it, it is your prompt set and your data | It lives in our platform |
| 12 | What is the cheapest thing I could do without you? | An immediate, specific, unpaid answer | Deflection back to the retainer |
Questions 4 and 5 are deliberate traps, and they are fair ones. Both have documented answers that contradict a large amount of what is sold as AEO. An agency that gets them wrong is either not reading primary sources or is hoping you have not.
For what it is worth, we are happy to be asked all twelve, and question 12 in particular. Our answer to that one is the free check itself, which comes before any proposal we would write.
What the work costs, and what drives the range
No credible benchmark survey exists for this yet, so treat what follows as our own observation of the market rather than published data. The useful part is not the numbers, it is understanding which variable you are paying for, because that is what makes two quotes differ by a factor of five.
| Scope | What it includes | What it does not | Who it suits |
|---|---|---|---|
| Diagnostic only | Access audit, entity check, baseline, prioritised findings | Any implementation | Anyone who does not yet know which problem they have |
| Technical remediation | Crawler policy, rendering, structured data, identity links | Content production | Sites where the constraint is engineering |
| Ongoing measurement | Frozen prompt set, repeated runs, reporting on citation share | Fixing what it finds | Teams who need to prove or disprove movement |
| Full programme | All of the above plus editorial and third-party work | Control of what other people publish | Companies with an entrenched competitor to displace |
Four things genuinely move the price. The number of engines you need measured, since each is a separate manual run. Whether your content requires engineering to become fetchable at all. How entrenched the incumbent in your category is, because displacing one is a different project from establishing presence. And whether you need somebody producing original data, which is the most expensive component and the one that produces citations nobody can copy.
What should not move the price is the number of pages on your site. Extractability is a property of the twenty paragraphs that matter, so per-page pricing on a three hundred page site is a large invoice attached to a small effect.
Red flags, specifically
Generic warnings are useless, so these are the ones we would actually walk away from.
A guaranteed position or citation rate. Nobody controls the retrieval systems, results vary run to run, and a guarantee in this field is either meaningless or is a claim somebody would struggle to substantiate. The FTC’s guidance on endorsements and objective claims turns on whether a claim can be backed with evidence at the time it is made, which is a reasonable standard to hold a proposal to even where it is not being enforced.
Proprietary schema, or an llms.txt file sold as the lever. Both are documented positions rather than matters of opinion. Google states there is no special structured data needed for its AI features, and llms.txt is a community proposal no provider guarantees anything about. Selling either as the mechanism is a reliable signal.
A visibility score with no denominator. If the number cannot be decomposed into how many prompts, on which engines, across how many runs, it cannot be checked, and a metric nobody can check is a reporting device rather than a measurement.
Measurement that lives in their platform. Your prompt set and your run history are the only continuity you have. If leaving means losing your baseline, you are renting the ability to know whether the work is working.
A content-volume proposal to a technical problem. Forty articles a month sounds like effort. If your logs show zero retrieval agent hits, all forty are invisible, and the proposal is answering a question you did not ask.
When you should not hire anyone
Three situations where an agency is the wrong purchase, and we have talked ourselves out of work in all three.
Your technical SEO is broken. If crawlers cannot reach or render your site, the fix is engineering that improves every surface at once, and paying an AEO premium for it is paying for the wrong label on the same work.
You have not run the free checks. Reading your robots.txt, counting agent hits in your logs and running ten prompts across four engines costs nothing and takes an afternoon. Do that first, because it changes what you are buying and it sometimes removes the need to buy.
Your problem is what the market says about you. If every engine independently reports the same weakness, your crawling is healthy and your entity resolves cleanly, then the systems are summarising accurate information. That is a product, pricing or positioning conversation, and no visibility programme addresses it. This is the one clients most dislike hearing and the one that saves the most money.
Building it in-house instead
Viable, and more often than agencies admit. The technical steps are one-off and documented. The measurement is manual but simple. What is genuinely hard to build internally is the comparative view, meaning knowing what normal looks like across a category, which is the one thing an outside party accumulates that you cannot.
A reasonable in-house setup is one engineer for a fortnight on access and rendering, one marketer owning a frozen prompt set with a recurring calendar slot, and an editor applying four extractability rules to new content as it is written. That covers most of the value. It does not cover original research or third-party corroboration, which are the parts that move recommendations.
The hybrid most of our better engagements settle into is a diagnostic and baseline from outside, then execution in-house, with an outside review each quarter. It is less revenue for us than a full retainer and it is what we would buy.
A scoring sheet for comparing two or three of them
Proposals are written to be hard to compare. Score them yourself on the things that predict whether the engagement works, rather than on how the deck looks. Two points where the answer is strong, one where it is adequate, zero where it is absent.
| Criterion | Why it predicts the outcome | Weight |
|---|---|---|
| Diagnosed before proposing | They read your logs and robots.txt before quoting | Double it |
| Named the binding constraint | One specific thing, not a programme of everything | Double it |
| Measurement you own and keep | Determines whether you can ever verify their work | Double it |
| Accurate on primary sources | Questions 4 and 5, which have documented answers | Single |
| Named practitioners on the call | Whether the people pitching are the people working | Single |
| Showed something that failed | Distinguishes experience from a case study library | Single |
| Named free work you should do first | The clearest signal of problem solving over packaging | Single |
| Told you a scenario they decline | An agency that fits everybody has no method | Single |
A proposal scoring well on the last five rows and zero on the first three is a good sales team attached to an unscoped project. That combination is the most common failure mode in this market, and it is the most expensive, because the work begins before anybody has established what the work is.
Score us on this sheet as well.
The three double-weighted rows are the ones we would want to be judged on, and they are also the reason we lead with a diagnostic instead of a proposal. Nobody can score the first two honestly without having read your logs first.
If you want to run the exercise properly, book an audit call and put us alongside two others. You will get the diagnosis, the baseline and a named constraint out of it whether or not you pick us.
Four things worth putting in the contract
None of these are unusual asks, and a reasonable supplier will agree to all four without friction. Reluctance on any of them is itself informative.
- The prompt set and run history are yours. Exportable in a plain format, on request, at any point. This is your only continuity if you change supplier.
- Reported figures carry their denominators. How many prompts, which engines, how many runs. A number without those is not reportable.
- Technical changes are documented as they are made. Anything touching robots.txt, headers or structured data, recorded with a date, so a future team can tell what changed and when.
- A defined exit that leaves you operational. A short handover producing the baseline, the current configuration and the outstanding list. Thirty days is plenty.
The third one matters more than it sounds. Most of the broken robots.txt files we find were changed deliberately by somebody who had a reason, left no record, and moved on. Documentation is the difference between a two-minute fix and a fortnight of archaeology.
The question to ask yourself before any of this
What decision would change if you had this information? It is worth answering honestly, because it determines whether you need a programme or a report.
If the answer is that you would fix whatever is broken, buy a diagnostic and act on it. If the answer is that you need to show a board that the company is visible in AI search, you are buying a reporting artefact, which is a legitimate purchase and a much smaller one. If the answer is that a competitor keeps appearing in answers and you do not, that is the strongest reason to spend, because you have a specific comparison to work against rather than a general anxiety.
The weakest reason, and the most common one we are approached with, is that AI search is clearly becoming important and somebody senior asked what the plan is. That is a real pressure and it is not a brief. Turning it into one means running the free checks first, because the findings decide whether this is an engineering ticket, an editorial habit or a genuine programme, and those differ by an order of magnitude in cost.
Frequently asked questions
Can any agency guarantee I will appear in ChatGPT?
No. Nobody operates the retrieval systems, answers vary between runs, and no provider offers placement. What can be promised is diagnostic work, specific technical remediation and honest measurement. A guarantee of position is a claim nobody could substantiate.
How long should a first engagement be?
Short enough to test the relationship on something verifiable. A diagnostic with a baseline is a few weeks and produces a document you keep regardless of what happens next. Signing a twelve-month programme before anybody has read your logs means committing to a scope nobody has scoped.
Should I hire an AEO specialist or my existing SEO agency?
Ask your current agency questions 3, 4 and 5 from the list above. If they answer well, they are the cheaper and better-integrated option, since most of the foundation is shared. If they cannot name the individual crawlers or do not know what Google publishes about AI features, you have your answer without changing supplier yet.
What deliverable should the first month produce?
A named binding constraint, a frozen prompt set you own, a recorded baseline across engines, and a prioritised list where the top item is the cheapest thing with the largest effect. If month one produces a content calendar instead, the diagnosis was skipped.
Is it too early to spend anything on this?
The free diagnostic work is never too early, because a crawler block costs you nothing to fix and everything to leave. Large programmes are easier to justify once you know the constraint is not a single line in a text file. Establishing a baseline early is the part with real time value, since you cannot measure a trend backwards.
What if the honest answer is that I am not going to win my category?
Then the useful goal changes rather than disappearing. Retrieval works on passages, so a narrower question you genuinely own is winnable even where the broad category question is not. Being the recommended answer for one specific buyer type beats being absent from a contest against an entrenched incumbent.
If you want the free checks first, the complete guide to AEO walks through the five gates and what to test at each one. Doing that before any sales call is the single best-value hour available to you.
The best outcome of this page is that you hire well, including if that means hiring nobody.
Take the twelve questions into every call you have booked, ours included, and score all of them on the same sheet. An agency that gets worse under specific questioning has told you something useful for free.
When you want a diagnosis rather than a proposal, book an audit call. We work the constraints in order, hand over the baseline as yours to keep, and say plainly when the honest answer is that you do not need an agency yet.
