AEO / GEO14 min read

Which agency is best for AEO and AI visibility reporting?

The best agency for AEO reporting is the one that can open a live dashboard showing prompts, engines, dates and citation counts for a current client. Not a slide, not a screenshot, a dashboard. Most agencies selling AEO in 2026 cannot do this, and asking to see it takes five minutes and eliminates about two thirds of a shortlist.

Roman Daneghyan - مؤلف مدونة في The Business Rover وSEO ووكالة النمو العضوي
تم التحديث في August 28, 2026

AI visibility has a measurement problem, and the measurement problem is why the vendor market is so hard to judge. There is no Search Console for ChatGPT. No official impressions feed from Perplexity. No rank tracker that Google endorses for AI Overviews. Every number you will ever see about your AI visibility was produced by somebody sampling prompts and counting what came back.

That makes methodology the whole product. Two agencies can report on the same brand in the same month and disagree wildly, purely because one sampled 30 prompts once and the other sampled 200 prompts weekly across four engines. Neither is lying. Only one is useful.

This page sets out what a credible reporting standard looks like, so you can grade proposals against something concrete instead of vibes. If the underlying concepts are new, start with what is AEO and how to measure AI visibility, then come back and use the tests below on your shortlist.

What does proper AEO reporting actually include?

Six components: a locked prompt set, multi-engine sampling, citation and mention counts, competitive share of voice, source-level attribution, and a written interpretation. Miss any one and the report becomes decorative.

The six components of a credible AEO report. The last column is what you get instead when an agency is improvising, which is the version most buyers are shown.

ComponentWhat it meansThe weak substitute
Locked prompt set50-300 buyer questions, defined once, sampled the same way every cycleAd-hoc prompts rewritten each month so trends cannot be compared
Multi-engine samplingChatGPT, Perplexity, Gemini, Claude and Google AI Overviews tracked separatelyOne ChatGPT screenshot presented as AI visibility
Citation and mention countsHow often you are named, and separately, how often you are linkedA single 'visibility score' with no stated formula
Share of voiceYour appearance rate against a fixed competitor set on the same promptsCompetitor names mentioned anecdotally in a call
Source attributionWhich URLs and third-party sources the models actually pulled fromAn assumption that your own site is the source
Written interpretationWhat moved, why, and what changes next month as a resultA dashboard link and no recommendation

Source attribution is the row most people underrate. Knowing that ChatGPT recommends three competitors and not you is a symptom. Knowing that it pulled all three from the same industry roundup you are absent from is a plan. The gap between those two reports is the difference between paying for monitoring and paying for optimisation.

Which AI visibility metrics actually matter?

Four: citation rate, share of voice, sentiment and position, and source concentration. Everything else is either a rollup of those or a number invented to fill a dashboard.

The two metrics that carry most of the weight

Citation rate = prompts where your brand appears / total prompts sampled
Share of voice = your appearances / (your appearances + tracked competitor appearances)

Both are calculated per engine and then rolled up, never averaged blindly across engines. A brand can hold 40% citation rate in Perplexity and 6% in ChatGPT, and the blended 23% would hide the only fact that matters: the engine your buyers actually use is the one ignoring you.

The four metrics worth arguing about in a monthly review, and what each one should trigger when it moves in the wrong direction.

MetricWhat it tells youWhat a decline should trigger
Citation rateWhether you are in the answer at allContent and entity work on the prompt clusters that dropped
Share of voiceWhether you are winning relative to named rivalsCompetitive teardown of the sources naming them instead of you
Sentiment and positionWhether being named is helping you or damaging youCorrective content and outreach to the sources shaping the description
Source concentrationHow dependent your visibility is on a few third-party pagesDiversification of the corroborating sources before one of them changes

Sentiment deserves more attention than it gets. Being cited as the cheap option, or being described with a feature you deprecated two years ago, is worse than not being cited at all. It is also fixable, but only if somebody is reading the answers rather than counting them.

If an agency reports a single blended AI visibility score without publishing the formula behind it, ask for the formula in writing. Composite scores are how weak sampling gets hidden inside a number that always seems to go up.

How is AEO reporting different from SEO reporting?

SEO reporting measures a stable, queryable index. AEO reporting measures a probabilistic system that can give two different answers to the same question five minutes apart. That single difference drives every methodological choice.

The structural differences between SEO and AEO measurement, and why AEO reports cost more to produce.

DimensionSEO reportingAEO reporting
Data sourceSearch Console, rank trackers, analyticsSampled model responses, no official feed
Unit of measurementPosition for a keywordPresence and framing inside a generated answer
StabilityReasonably stable day to dayVolatile; repeat sampling is mandatory
Traffic linkageDirect - clicks are countedPartial - many answers produce no click at all
Competitive viewWho ranks above youWho is named alongside you, and how
Cost driverTool licencesSampling volume across engines and prompts

The traffic linkage row is the uncomfortable one. A buyer can read your entire positioning inside an AI answer, form an opinion, and arrive three weeks later as direct traffic. Your analytics will call that a brand visit. Any agency claiming precise revenue attribution from AI answers is overselling, and the good ones say so before you ask.

How do you test whether an agency can really measure ChatGPT visibility?

Give them one prompt and ten minutes. Ask them to show you, live, how your brand performs on a question your buyers actually ask, across at least two engines, with an explanation of how they would track it over time.

  • Ask them to open a current client's dashboard and filter to a single prompt cluster
  • Ask how many prompts that client is tracked on, and how often those prompts are sampled
  • Ask what they do when the same prompt returns different citations on consecutive runs
  • Ask which engines they cannot measure well, and why
  • Ask them to export a raw prompt-and-response pair, not a summary

The fourth question is the best of the five. Every honest practitioner has a list of surfaces that are hard to sample reliably, and personalised or logged-in experiences are near the top of it. An agency that claims full coverage of everything, including whatever a specific user sees in their own ChatGPT account, is describing something that does not exist.

Tooling is a fair signal too. Purpose-built platforms such as PromptRush exist to map which prompts trigger competitor mentions and where the citations come from, and an agency running that kind of sampling can tell you where you stand within a week. One relying on manual spot checks will send you a schema audit and describe it as AEO. Our own standard for this is documented under AEO reporting.

Show me the dashboard is the fastest disqualifier in agency selection right now. Nobody who does this work needs a week to prepare it.

- Roman Daneghyan, The Business Rover

What tools do agencies use for AI visibility tracking?

Three categories: prompt-tracking platforms, general SEO suites with an AI module bolted on, and in-house sampling built with model APIs. Each has a different failure mode, and the choice tells you how seriously the agency takes the discipline.

How AI visibility tooling breaks down in 2026, including what each approach tends to get wrong.

ApproachStrengthWhere it falls short
Prompt-tracking platforms (for example PromptRush)Purpose-built prompt sets, competitor mapping, repeat sampling across enginesCoverage varies by engine and language; still needs someone to interpret it
SEO suites with AI modulesConvenient if you already pay for the suite; ties into keyword dataShallow sampling frequency and small prompt allowances
In-house API samplingFull control of prompts, frequency and storageModel APIs do not always match the consumer product; engineering cost is real

The API caveat matters more than most agencies admit. What the ChatGPT API returns is not always what a person sees in the ChatGPT app, because the consumer product layers on retrieval, memory and product-specific behaviour. A serious measurement program acknowledges the gap and samples the consumer surfaces too, rather than pretending the API is a clean proxy.

How often should prompts be sampled?

Weekly at minimum, daily for high-value prompt clusters. Monthly sampling produces a number you cannot trust, because a single run of a volatile system is a snapshot of noise.

Model responses vary run to run. Ask the same shortlist question three times and you can get three different sets of names. The fix is repetition: sample each prompt multiple times per cycle and report the appearance rate across runs rather than a binary yes or no. Any agency that reports 'we appear for this prompt' without saying how many runs that was based on has skipped the only step that makes the claim meaningful.

Sizing a prompt set that will actually detect movement

Prompts = buyer questions per stage x stages in your funnel x variants per question
Runs per cycle = 3 to 5 per prompt per engine

For most B2B companies this lands between 80 and 250 prompts. Fewer than 50 and normal volatility swamps any real change. More than 300 and the sampling cost rises faster than the insight, unless you are running multiple markets or languages.

What does a good monthly AEO report look like?

One page of numbers, one page of interpretation, and a decision at the end. If the report does not change what happens next month, the agency is charging you for monitoring.

  • Headline: citation rate and share of voice per engine, with the change from last cycle and the sample size behind it
  • Movement: which prompt clusters gained or lost, and the most plausible cause
  • Competitive: which rival gained the most, and which sources are naming them
  • Sources: the third-party pages the models pulled from, ranked by how often they appeared
  • Actions: what shipped last month, what ships next month, and what got dropped

The sources section is where reporting turns into strategy. When the same three industry roundups keep feeding an engine's answers, the work becomes getting represented in those roundups, which is closer to digital PR than to on-page optimisation. That is why serious AEO programs usually carry a link building and digital PR component rather than treating content as the only lever.

What are the warning signs in AEO reporting?

Five, and each one has a version that sounds impressive on a sales call. Any single one is a question. Three or more means the agency is selling SEO with new vocabulary.

Warning signs in AEO reporting, what they usually mean, and the question that exposes them.

Warning signWhat it usually meansAsk this
A single proprietary visibility scoreThe formula hides thin samplingWhat is the formula, and what sample size sits behind it
Screenshots without dates or promptsCherry-picked runs from a volatile systemCan you export the raw prompt and response pairs
One engine onlyThey can measure ChatGPT and are guessing at the restShow me the same prompt cluster in Perplexity and Gemini
Guaranteed citationsThey do not understand or are misrepresenting retrievalWhat exactly are you guaranteeing in the contract
Traffic charts labelled as AI visibilityStandard SEO reporting with a new cover slideWhich of these sessions came from an AI surface, and how do you know

Guaranteed citations is the 2026 version of guaranteed number one rankings. Nobody controls how a model retrieves and synthesises sources. An agency can commit to process, cadence, and the work it will ship. It cannot commit to what ChatGPT will say about you in November.

How much does AEO reporting cost?

Reporting alone runs $1,000 to $4,000 a month depending on prompt volume, engines and markets. Bundled into a full program it typically adds 15% to 25% on top of an SEO retainer.

The cost scales with sampling, not with the dashboard. Two hundred prompts sampled five times across four engines is four thousand model calls a cycle before anyone reads a single answer. That is why the cheap version of this always turns out to be thirty prompts checked once a month, and why the resulting trend line is unusable.

Whether that spend makes sense depends on how much of your category's buying research has already moved to AI assistants. In software, professional services and anything with a considered purchase, it has moved a lot. Our delivery model for it is described under AEO services, and the specialist end of the vendor market is mapped in the AEO and GEO agency roundup.

How do you connect AI visibility to revenue?

Imperfectly, and anyone who tells you otherwise is selling. What you can do is measure the correlation between citation gains and downstream demand, and use self-reported attribution to fill the gap analytics cannot.

A defensible way to value AI visibility

Estimated influenced pipeline = prompt volume x citation rate x assisted conversion rate x average deal value
Confidence check = share of new deals citing AI assistants in your 'how did you hear about us' field

Prompt volume is estimated from search demand for the equivalent queries, so treat it as directional rather than precise. The confidence check is the honest half of the equation: if nobody in your sales pipeline mentions AI assistants, no dashboard number justifies the spend yet.

Add a free-text source field to your demo or contact form and read it monthly. It is unglamorous and it works. Companies that added one in 2025 are the ones that can now say what proportion of their pipeline started inside an AI conversation, and they are making better budget decisions than the ones staring at attribution models that were never designed for zero-click research.

So which agency should you pick for AEO reporting?

The one that shows you the dashboard in the first call, tells you which engines they cannot measure well, and connects every metric to something they will do differently next month. That combination is still rare, which is exactly why it is a useful filter.

Run the five-minute test on your shortlist before you compare prices. The wider selection checklist is in how to choose an SEO agency, the category framework by business model is in the best SEO agency for the US market, and if you are still working out what an agency does at all, start with what is an SEO agency. Everything else lives in the answers library.

If you want to see what your current AI visibility looks like before you hire anyone, we run the baseline as a standalone exercise. Ask for a proposal and you will get the prompt set and the numbers, whether or not you go on to work with us. Teams that need the wider AI program rather than reporting alone should look at our AI SEO agency page.

Frequently asked questions

What is AEO reporting?
AEO reporting measures how often and how favourably your brand appears in AI-generated answers. A credible report includes a locked prompt set, sampling across ChatGPT, Perplexity, Gemini and AI Overviews, citation rate, competitive share of voice, the third-party sources the models pulled from, and a written interpretation of what changed.
Which metrics matter most for AI visibility?
Citation rate, share of voice against a fixed competitor set, sentiment and position within the answer, and source concentration. Each should be reported per engine rather than blended, because a strong Perplexity presence can hide a total absence from the engine your buyers actually use.
How do I know if an agency can really measure ChatGPT visibility?
Ask them to open a live client dashboard filtered to one prompt cluster, tell you the sample size and frequency, and export a raw prompt and response pair. Then ask which engines they cannot measure reliably. Anyone claiming complete coverage of every surface, including personalised sessions, is overstating what is possible.
How often should AI visibility be tracked?
Weekly at minimum, with three to five runs per prompt per engine, and daily for high-value prompt clusters. Model answers vary between runs, so single monthly checks report noise. Appearance rate across repeated runs is the only version of the number worth acting on.
Can an agency guarantee citations in ChatGPT?
No. No agency controls how a model retrieves or synthesises sources, and guaranteed citations is the 2026 equivalent of guaranteed number one rankings. Agencies can commit to sampling cadence, shipped work and reporting standards, which is what a real contract should specify.
How much does AEO reporting cost?
Standalone reporting runs roughly $1,000 to $4,000 a month depending on prompt volume, number of engines and markets covered. Inside a full program it usually adds 15% to 25% to an SEO retainer, because cost scales with sampling volume rather than with the dashboard itself.
Can you attribute revenue to AI search visibility?
Not precisely. Many AI answers produce no click, so buyers arrive later as direct or branded traffic. The workable approach is to track correlation between citation gains and demand, and add a free-text 'how did you hear about us' field so prospects can tell you directly.