Worked example

What a monthly AEO report actually looks like.

The short answer

A monthly AEO report is four tables and a method note, reported per engine and traceable to the answers behind every number. The four are a per-engine visibility grid, a share-of-voice table against named competitors, a citation-gap matrix and the source list. Anything that arrives as a single score with no method note is not a report — it is a screenshot.

About the numbers on this page

Every figure in the example tables below is illustrative. It is constructed to show the shape and the arithmetic of a credible report, using published benchmark ranges as plausible values. It is not client data, it is not a real brand's measurement, and no reader should treat it as a benchmark for their own category. The sourced statistics are marked separately and link to their studies.

Why the report is built this way

Engines disagree, and a single run is not a measurement.

The report structure below is not a stylistic preference. It follows from three published findings: the engines cite largely different webs, they cite at wildly different depths, and a single sample of a probabilistic system tells you almost nothing.

69.6%Share of cited domains that were cited by only one engine, across 127,198 citations and ~16,400 answers from five engines — which is why a blended cross-engine score hides where you actually appear (SurfacedBy, Mar–Jun 2026)SurfacedBy, AI citation study: engine overlap
3.7 vs 11.0Average sources cited per answer by ChatGPT versus Gemini in the same study — the two engines are not competing for the same slots, so one row per engine is the minimum honest resolutionSurfacedBy, AI citation study: engine overlap
85%Share of AI-answer citations originating from third-party sources rather than a brand's own domain — so the source list is a required section, not an appendix (AirOps)Omnibound, AI Search Statistics

The document

A monthly report answers four questions, in this order.

A monthly Answer Engine Optimization report exists to answer four questions for the person paying for the work: where do we appear, how do we compare to the competitors we care about, what is standing between us and the answers we are missing, and which sources are feeding the engines that already name us. Everything else in the document is either evidence for one of those four answers or decoration. The order matters, because each table is only interpretable once the one before it has been read.

Before any of them comes the method note. It is one short paragraph, it goes on page one, and it states the prompt set size, the number of runs per prompt, the engines sampled, the geography and account state, and the date range. Without it, every number that follows is unfalsifiable. A report whose method note is missing is not a rigorous report with an omission — it is a document that cannot be checked, which is a different category of thing. The section-by-section acceptance criteria for each table are set out in our guide to what a good AI visibility report contains.

Table 1: the per-engine visibility grid

The first table is one row per engine and one column per outcome, never blended into a single number. The three outcomes are genuinely different events: a citation is a link or named source in the answer, a mention is the brand named without being sourced, and a recommendation is the engine actively proposing the brand as an option. A brand can be cited constantly and recommended never, which is a content problem; it can be recommended without being cited, which lives in the model's parametric memory rather than in retrieval. A report that collapses the three loses the information needed to tell those apart.

Blending engines into one figure is not merely imprecise. In the SurfacedBy study of 127,198 citations, 69.6% of cited domains were cited by only one engine and just 2.7% were cited by all five, and the average number of sources per answer ran from 3.7 in ChatGPT to 11.0 in Gemini. Averaging across engines with that spread produces a number that describes no engine at all, and worse, it can move for reasons that have nothing to do with the work — a shift in the engine mix of the sample will move a blended score on its own.

Table 2: share of voice against named competitors

Share of voice is only meaningful against a fixed, named comparison set agreed with the client before the month starts. The competitor list must not change between months without the change being flagged in the method note, because quietly dropping the competitor who beat you is the easiest way to manufacture improvement. The table shows, for each competitor and for the brand, the share of answers in which each was named — the sum will not be 100%, because most answers name several brands. The honest version also reports the share of answers that named nobody, which in many categories is the largest cell and is genuinely good news: it is the addressable space.

Table 3: the citation-gap matrix

This is the table that turns measurement into work. For each prompt where a competitor was named and the brand was not, it records which engine, which competitor, and which source the engine cited to justify naming them. That third column is the whole point, because roughly 85% of AI-answer citations come from third-party sources rather than the brand's own domain. The gap is therefore almost never solved on the brand's own website; it is solved wherever the engine went looking.

A recent measurement framework sharpens this. Analysing 602 controlled prompts, 21,143 search-layer citations and 18,151 fetched pages across ChatGPT, Google AI Overview/Gemini and Perplexity, the authors separate citation selection — being retrieved into the source list — from citation absorption, where a cited page actually contributes language, evidence or structure to the answer. A page can be selected constantly and absorbed rarely. Higher-absorption pages were longer, more structured, more semantically aligned with the query and richer in extractable evidence such as definitions, statistics and comparison tables. A gap matrix that only counts appearances in source lists measures the cheaper half of the problem.

Table 4: the source list

The fourth table lists every URL the engines cited in answers where the brand appeared, grouped by domain with a count. It is the least glamorous section and often the most actionable, because it names the specific third-party surfaces doing the work — a directory profile, a review page, a comparison article, an industry publication. Over several months it becomes a map of the brand's real source graph and the evidence base for deciding where next month's off-site effort should go.

What the report should not contain

Three things routinely appear in monthly AEO reports and should not. A single blended visibility score with no published formula: if the reader cannot reconstruct it, it cannot be audited. A website traffic graph presented as proof of AI visibility: AI-driven visits largely arrive without a readable referrer — one analysis of 446,405 visits found 70.6% landing in the 'Direct' bucket — so traffic is a poor, lagging proxy for a channel measured properly by citation, mention and recommendation rate. And week-over-week movement on a stochastic system: below about ten runs per prompt, a small change is noise, and reporting it as progress is reporting noise.

Illustrative — Table 1

The per-engine visibility grid, as it would appear in the document.

An example of the shape and arithmetic only. These numbers are constructed for illustration and are not a benchmark, a client result or a real brand's measurement. Note that the engines disagree enough that no single average would describe any row.

Illustrative example. Method note that would accompany it: 60-prompt portfolio, 10 runs per prompt per engine, five engines, US English, logged-out, 1–30 of the reporting month. Month-over-month change would be reported per cell as an interval, never as a single point.
EngineCitation rateMention rateRecommendation rate
ChatGPT8%21%12%
Google AI Mode19%26%9%
Perplexity24%28%14%
Gemini22%25%7%
Claude11%18%10%
Blended average (do not report)16.8%23.6%10.4%

How to produce it

Building the month's report, step by step.

The same sequence every month. The discipline is that the prompt set and the competitor set are fixed in advance and any change to either is declared, because a moving denominator can produce improvement out of nothing.

Freeze the prompt portfolio and the competitor set

Agree a fixed set of buyer-intent prompts and a named competitor list before the month starts, and write both into the method note. If a prompt is added or retired, say so and show the affected numbers both ways for one month.

Sample repeatedly, per engine, never blended

Run each prompt at least ten times on each engine before treating a rate as stable, holding geography and account state constant. A 50-prompt set run ten times says more about stability than a 500-prompt set run once.

Record three outcomes separately

For every answer, log whether the brand was cited, mentioned, or recommended. These are different events with different causes and different fixes; one combined 'visibility' figure destroys the information needed to choose next month's work.

Capture the source URL behind every appearance

Log the exact URLs the engine cited, not just whether it cited something. This is what makes the source list possible, and the only way to see which third-party surfaces feed the answers, given that roughly 85% of citations are off-site.

Build the gap matrix from the misses, not the wins

Filter to the prompts where a competitor was named and the brand was not, then record the engine, the competitor and the source that justified them. Note where possible whether the rival page was merely listed or visibly shaped the wording.

Report the change with its uncertainty

Show month-over-month movement as an interval, not a point estimate, and say plainly when a change sits inside the sampling noise. A report willing to say 'this is not distinguishable from run-to-run variance' is the more credible one.

Close with actions tied to specific rows

Every recommended action should point at a numbered row in the gap matrix or the source list. An action list that cannot be traced to a measured gap is a content calendar wearing a report's clothing.

Reading the example honestly

Three things the illustrative tables above do not mean.

Worked examples are useful for showing structure and dangerous when mistaken for data. These are the three misreadings to avoid.

The percentages are not benchmarks

They were constructed to demonstrate arithmetic and shape. Real rates vary enormously by category, prompt phrasing, geography and how narrowly the brand's market is defined, and there is no published figure that would make a cross-category benchmark meaningful.

The engine ordering is not a ranking

The illustrative rows are not a claim that Perplexity cites any given brand more than ChatGPT does. Published data shows the engines differ sharply in how many sources they cite per answer and which domains they draw on, but the direction for any specific brand is an empirical question its own measurement has to answer.

A good report does not promise an outcome

Measuring the channel well does not make an engine's answer controllable. Answers are probabilistic and change with every model update, so a report's job is to describe what happened and narrow the evidence gap — never to underwrite a placement.

Definition

Monthly AEO report

Monthly AEO report

A recurring document that measures how AI engines cited, mentioned and recommended a brand over a month, reported per engine against a fixed prompt portfolio and traceable to the exact answers and sources behind each figure.

A monthly AEO report differs from a marketing performance report in its unit of observation. The unit is not a session or a ranking position but an answer: a single generated response to a prompt, from one engine, at one moment, citing a specific set of sources. Because the same prompt produces different answers across runs, the report is a sampling exercise, and its credibility rests on the sampling design being declared rather than on the numbers being flattering. The minimum viable version is a method note, a per-engine grid of citation, mention and recommendation rate, a share-of-voice table against a fixed competitor set, a citation-gap matrix and a source list. It describes a probabilistic system; it cannot promise a placement in one.

Disclosure

Why this portal publishes a report template and not a report service.

This site is an independent reference portal on AEO. It does not sell audits, reports or optimization work, and it does not broker introductions or take a fee from the agencies it lists. Companies that want this report produced for them should choose a provider from the directory of recommended agencies and use the buyer's checklist to test the method before signing.

The portal's operator also runs the agency Blobic, which appears in the directory under exactly the same published criteria as every other listing, carrying a disclosure badge. Blobic is not ranked above anyone, no placement in the directory is or has ever been for sale, and the inclusion criteria are published in full so any reader can check the entries against them. We state the relationship every time the text touches the directory's neutrality, because the directory's credibility is the only asset the portal has.

FAQ

Common questions about monthly AEO reporting.

What should an AEO report include?

A method note stating the prompt set, runs per prompt, engines, geography and date range; a per-engine grid of citation, mention and recommendation rate; a share-of-voice table against a fixed, named competitor set; a citation-gap matrix listing the prompts where competitors were named and you were not, with the source the engine cited; and a source list of the URLs feeding the answers that named you. Actions at the end should point at specific rows in those tables.

How often should an AI visibility report be produced?

Monthly is the practical cadence for reporting, with sampling happening more frequently underneath it. Answers are probabilistic and vary run to run, so a shorter reporting window tends to surface sampling noise as if it were progress. A rate should be based on at least ten runs per prompt per engine before month-over-month movement is worth interpreting.

Why should a report never blend the engines into one score?

Because the engines cite largely different webs. Across 127,198 citations from five engines, 69.6% of cited domains were cited by only one engine and just 2.7% by all five, and average sources per answer ranged from 3.7 in ChatGPT to 11.0 in Gemini. An average across engines with that spread describes none of them, and it can move purely because the engine mix of the sample changed.

Is website traffic a valid measure of AI visibility?

No, it is a weak and lagging proxy. Most AI-driven visits arrive without a readable referrer — one analysis of 446,405 visits found 70.6% landing in the 'Direct' bucket in GA4 — so traffic under-counts the channel and cannot tell you which prompt, engine or answer produced a visit. Measure the channel by citation, mention and recommendation rate against a fixed prompt portfolio, and treat traffic as a secondary signal.

Can an agency guarantee the numbers in the report will improve?

No. AI answers are generated probabilistically and are re-synthesized with every model update, so no provider can guarantee a citation, a mention or a recommendation. What an agency can commit to is method: a declared prompt portfolio, honest sampling, per-engine reporting and a traceable evidence gap. A guaranteed placement is the clearest red flag in the category.

Are the example numbers on this page real?

No. Every figure in the illustrative tables is constructed to demonstrate the structure and arithmetic of a report. They are not client data, not a real brand's measurement and not a benchmark for any category. The sourced statistics are marked separately and link to the published studies behind them.

Next

Hold a report to the template — or apply to be listed.

Companies comparing providers can start from the directory and the inclusion criteria. Agencies that already report this way can propose themselves for the directory.