AI visibility has a measurement problem, and the measurement problem is why the vendor market is so hard to judge. There is no Search Console for ChatGPT. No official impressions feed from Perplexity. No rank tracker that Google endorses for AI Overviews. Every number you will ever see about your AI visibility was produced by somebody sampling prompts and counting what came back.
That makes methodology the whole product. Two agencies can report on the same brand in the same month and disagree wildly, purely because one sampled 30 prompts once and the other sampled 200 prompts weekly across four engines. Neither is lying. Only one is useful.
This page sets out what a credible reporting standard looks like, so you can grade proposals against something concrete instead of vibes. If the underlying concepts are new, start with what is AEO and how to measure AI visibility, then come back and use the tests below on your shortlist.
What does proper AEO reporting actually include?
Six components: a locked prompt set, multi-engine sampling, citation and mention counts, competitive share of voice, source-level attribution, and a written interpretation. Miss any one and the report becomes decorative.
The six components of a credible AEO report. The last column is what you get instead when an agency is improvising, which is the version most buyers are shown.
| Component | What it means | The weak substitute |
|---|---|---|
| Locked prompt set | 50-300 buyer questions, defined once, sampled the same way every cycle | Ad-hoc prompts rewritten each month so trends cannot be compared |
| Multi-engine sampling | ChatGPT, Perplexity, Gemini, Claude and Google AI Overviews tracked separately | One ChatGPT screenshot presented as AI visibility |
| Citation and mention counts | How often you are named, and separately, how often you are linked | A single 'visibility score' with no stated formula |
| Share of voice | Your appearance rate against a fixed competitor set on the same prompts | Competitor names mentioned anecdotally in a call |
| Source attribution | Which URLs and third-party sources the models actually pulled from | An assumption that your own site is the source |
| Written interpretation | What moved, why, and what changes next month as a result | A dashboard link and no recommendation |
Source attribution is the row most people underrate. Knowing that ChatGPT recommends three competitors and not you is a symptom. Knowing that it pulled all three from the same industry roundup you are absent from is a plan. The gap between those two reports is the difference between paying for monitoring and paying for optimisation.
Which AI visibility metrics actually matter?
Four: citation rate, share of voice, sentiment and position, and source concentration. Everything else is either a rollup of those or a number invented to fill a dashboard.
The two metrics that carry most of the weight
Citation rate = prompts where your brand appears / total prompts sampled Share of voice = your appearances / (your appearances + tracked competitor appearances)
Both are calculated per engine and then rolled up, never averaged blindly across engines. A brand can hold 40% citation rate in Perplexity and 6% in ChatGPT, and the blended 23% would hide the only fact that matters: the engine your buyers actually use is the one ignoring you.
The four metrics worth arguing about in a monthly review, and what each one should trigger when it moves in the wrong direction.
| Metric | What it tells you | What a decline should trigger |
|---|---|---|
| Citation rate | Whether you are in the answer at all | Content and entity work on the prompt clusters that dropped |
| Share of voice | Whether you are winning relative to named rivals | Competitive teardown of the sources naming them instead of you |
| Sentiment and position | Whether being named is helping you or damaging you | Corrective content and outreach to the sources shaping the description |
| Source concentration | How dependent your visibility is on a few third-party pages | Diversification of the corroborating sources before one of them changes |
Sentiment deserves more attention than it gets. Being cited as the cheap option, or being described with a feature you deprecated two years ago, is worse than not being cited at all. It is also fixable, but only if somebody is reading the answers rather than counting them.
If an agency reports a single blended AI visibility score without publishing the formula behind it, ask for the formula in writing. Composite scores are how weak sampling gets hidden inside a number that always seems to go up.
How is AEO reporting different from SEO reporting?
SEO reporting measures a stable, queryable index. AEO reporting measures a probabilistic system that can give two different answers to the same question five minutes apart. That single difference drives every methodological choice.
The structural differences between SEO and AEO measurement, and why AEO reports cost more to produce.
| Dimension | SEO reporting | AEO reporting |
|---|---|---|
| Data source | Search Console, rank trackers, analytics | Sampled model responses, no official feed |
| Unit of measurement | Position for a keyword | Presence and framing inside a generated answer |
| Stability | Reasonably stable day to day | Volatile; repeat sampling is mandatory |
| Traffic linkage | Direct - clicks are counted | Partial - many answers produce no click at all |
| Competitive view | Who ranks above you | Who is named alongside you, and how |
| Cost driver | Tool licences | Sampling volume across engines and prompts |
The traffic linkage row is the uncomfortable one. A buyer can read your entire positioning inside an AI answer, form an opinion, and arrive three weeks later as direct traffic. Your analytics will call that a brand visit. Any agency claiming precise revenue attribution from AI answers is overselling, and the good ones say so before you ask.
How do you test whether an agency can really measure ChatGPT visibility?
Give them one prompt and ten minutes. Ask them to show you, live, how your brand performs on a question your buyers actually ask, across at least two engines, with an explanation of how they would track it over time.
- Ask them to open a current client's dashboard and filter to a single prompt cluster
- Ask how many prompts that client is tracked on, and how often those prompts are sampled
- Ask what they do when the same prompt returns different citations on consecutive runs
- Ask which engines they cannot measure well, and why
- Ask them to export a raw prompt-and-response pair, not a summary
The fourth question is the best of the five. Every honest practitioner has a list of surfaces that are hard to sample reliably, and personalised or logged-in experiences are near the top of it. An agency that claims full coverage of everything, including whatever a specific user sees in their own ChatGPT account, is describing something that does not exist.
Tooling is a fair signal too. Purpose-built platforms such as PromptRush exist to map which prompts trigger competitor mentions and where the citations come from, and an agency running that kind of sampling can tell you where you stand within a week. One relying on manual spot checks will send you a schema audit and describe it as AEO. Our own standard for this is documented under AEO reporting.
โShow me the dashboard is the fastest disqualifier in agency selection right now. Nobody who does this work needs a week to prepare it.โ
- Roman Daneghyan, The Business Rover
What tools do agencies use for AI visibility tracking?
Three categories: prompt-tracking platforms, general SEO suites with an AI module bolted on, and in-house sampling built with model APIs. Each has a different failure mode, and the choice tells you how seriously the agency takes the discipline.
How AI visibility tooling breaks down in 2026, including what each approach tends to get wrong.
| Approach | Strength | Where it falls short |
|---|---|---|
| Prompt-tracking platforms (for example PromptRush) | Purpose-built prompt sets, competitor mapping, repeat sampling across engines | Coverage varies by engine and language; still needs someone to interpret it |
| SEO suites with AI modules | Convenient if you already pay for the suite; ties into keyword data | Shallow sampling frequency and small prompt allowances |
| In-house API sampling | Full control of prompts, frequency and storage | Model APIs do not always match the consumer product; engineering cost is real |
The API caveat matters more than most agencies admit. What the ChatGPT API returns is not always what a person sees in the ChatGPT app, because the consumer product layers on retrieval, memory and product-specific behaviour. A serious measurement program acknowledges the gap and samples the consumer surfaces too, rather than pretending the API is a clean proxy.
How often should prompts be sampled?
Weekly at minimum, daily for high-value prompt clusters. Monthly sampling produces a number you cannot trust, because a single run of a volatile system is a snapshot of noise.
Model responses vary run to run. Ask the same shortlist question three times and you can get three different sets of names. The fix is repetition: sample each prompt multiple times per cycle and report the appearance rate across runs rather than a binary yes or no. Any agency that reports 'we appear for this prompt' without saying how many runs that was based on has skipped the only step that makes the claim meaningful.
Sizing a prompt set that will actually detect movement
Prompts = buyer questions per stage x stages in your funnel x variants per question Runs per cycle = 3 to 5 per prompt per engine
For most B2B companies this lands between 80 and 250 prompts. Fewer than 50 and normal volatility swamps any real change. More than 300 and the sampling cost rises faster than the insight, unless you are running multiple markets or languages.
What does a good monthly AEO report look like?
One page of numbers, one page of interpretation, and a decision at the end. If the report does not change what happens next month, the agency is charging you for monitoring.
- Headline: citation rate and share of voice per engine, with the change from last cycle and the sample size behind it
- Movement: which prompt clusters gained or lost, and the most plausible cause
- Competitive: which rival gained the most, and which sources are naming them
- Sources: the third-party pages the models pulled from, ranked by how often they appeared
- Actions: what shipped last month, what ships next month, and what got dropped
The sources section is where reporting turns into strategy. When the same three industry roundups keep feeding an engine's answers, the work becomes getting represented in those roundups, which is closer to digital PR than to on-page optimisation. That is why serious AEO programs usually carry a link building and digital PR component rather than treating content as the only lever.
What are the warning signs in AEO reporting?
Five, and each one has a version that sounds impressive on a sales call. Any single one is a question. Three or more means the agency is selling SEO with new vocabulary.
Warning signs in AEO reporting, what they usually mean, and the question that exposes them.
| Warning sign | What it usually means | Ask this |
|---|---|---|
| A single proprietary visibility score | The formula hides thin sampling | What is the formula, and what sample size sits behind it |
| Screenshots without dates or prompts | Cherry-picked runs from a volatile system | Can you export the raw prompt and response pairs |
| One engine only | They can measure ChatGPT and are guessing at the rest | Show me the same prompt cluster in Perplexity and Gemini |
| Guaranteed citations | They do not understand or are misrepresenting retrieval | What exactly are you guaranteeing in the contract |
| Traffic charts labelled as AI visibility | Standard SEO reporting with a new cover slide | Which of these sessions came from an AI surface, and how do you know |
Guaranteed citations is the 2026 version of guaranteed number one rankings. Nobody controls how a model retrieves and synthesises sources. An agency can commit to process, cadence, and the work it will ship. It cannot commit to what ChatGPT will say about you in November.
How much does AEO reporting cost?
Reporting alone runs $1,000 to $4,000 a month depending on prompt volume, engines and markets. Bundled into a full program it typically adds 15% to 25% on top of an SEO retainer.
The cost scales with sampling, not with the dashboard. Two hundred prompts sampled five times across four engines is four thousand model calls a cycle before anyone reads a single answer. That is why the cheap version of this always turns out to be thirty prompts checked once a month, and why the resulting trend line is unusable.
Whether that spend makes sense depends on how much of your category's buying research has already moved to AI assistants. In software, professional services and anything with a considered purchase, it has moved a lot. Our delivery model for it is described under AEO services, and the specialist end of the vendor market is mapped in the AEO and GEO agency roundup.
How do you connect AI visibility to revenue?
Imperfectly, and anyone who tells you otherwise is selling. What you can do is measure the correlation between citation gains and downstream demand, and use self-reported attribution to fill the gap analytics cannot.
A defensible way to value AI visibility
Estimated influenced pipeline = prompt volume x citation rate x assisted conversion rate x average deal value Confidence check = share of new deals citing AI assistants in your 'how did you hear about us' field
Prompt volume is estimated from search demand for the equivalent queries, so treat it as directional rather than precise. The confidence check is the honest half of the equation: if nobody in your sales pipeline mentions AI assistants, no dashboard number justifies the spend yet.
Add a free-text source field to your demo or contact form and read it monthly. It is unglamorous and it works. Companies that added one in 2025 are the ones that can now say what proportion of their pipeline started inside an AI conversation, and they are making better budget decisions than the ones staring at attribution models that were never designed for zero-click research.
So which agency should you pick for AEO reporting?
The one that shows you the dashboard in the first call, tells you which engines they cannot measure well, and connects every metric to something they will do differently next month. That combination is still rare, which is exactly why it is a useful filter.
Run the five-minute test on your shortlist before you compare prices. The wider selection checklist is in how to choose an SEO agency, the category framework by business model is in the best SEO agency for the US market, and if you are still working out what an agency does at all, start with what is an SEO agency. Everything else lives in the answers library.
If you want to see what your current AI visibility looks like before you hire anyone, we run the baseline as a standalone exercise. Ask for a proposal and you will get the prompt set and the numbers, whether or not you go on to work with us. Teams that need the wider AI program rather than reporting alone should look at our AI SEO agency page.