A monthly Answer Engine Optimization report exists to answer four questions for the person paying for the work: where do we appear, how do we compare to the competitors we care about, what is standing between us and the answers we are missing, and which sources are feeding the engines that already name us. Everything else in the document is either evidence for one of those four answers or decoration. The order matters, because each table is only interpretable once the one before it has been read.
Before any of them comes the method note. It is one short paragraph, it goes on page one, and it states the prompt set size, the number of runs per prompt, the engines sampled, the geography and account state, and the date range. Without it, every number that follows is unfalsifiable. A report whose method note is missing is not a rigorous report with an omission — it is a document that cannot be checked, which is a different category of thing. The section-by-section acceptance criteria for each table are set out in our guide to what a good AI visibility report contains.
Table 1: the per-engine visibility grid
The first table is one row per engine and one column per outcome, never blended into a single number. The three outcomes are genuinely different events: a citation is a link or named source in the answer, a mention is the brand named without being sourced, and a recommendation is the engine actively proposing the brand as an option. A brand can be cited constantly and recommended never, which is a content problem; it can be recommended without being cited, which lives in the model's parametric memory rather than in retrieval. A report that collapses the three loses the information needed to tell those apart.
Blending engines into one figure is not merely imprecise. In the SurfacedBy study of 127,198 citations, 69.6% of cited domains were cited by only one engine and just 2.7% were cited by all five, and the average number of sources per answer ran from 3.7 in ChatGPT to 11.0 in Gemini. Averaging across engines with that spread produces a number that describes no engine at all, and worse, it can move for reasons that have nothing to do with the work — a shift in the engine mix of the sample will move a blended score on its own.
Table 2: share of voice against named competitors
Share of voice is only meaningful against a fixed, named comparison set agreed with the client before the month starts. The competitor list must not change between months without the change being flagged in the method note, because quietly dropping the competitor who beat you is the easiest way to manufacture improvement. The table shows, for each competitor and for the brand, the share of answers in which each was named — the sum will not be 100%, because most answers name several brands. The honest version also reports the share of answers that named nobody, which in many categories is the largest cell and is genuinely good news: it is the addressable space.
Table 3: the citation-gap matrix
This is the table that turns measurement into work. For each prompt where a competitor was named and the brand was not, it records which engine, which competitor, and which source the engine cited to justify naming them. That third column is the whole point, because roughly 85% of AI-answer citations come from third-party sources rather than the brand's own domain. The gap is therefore almost never solved on the brand's own website; it is solved wherever the engine went looking.
A recent measurement framework sharpens this. Analysing 602 controlled prompts, 21,143 search-layer citations and 18,151 fetched pages across ChatGPT, Google AI Overview/Gemini and Perplexity, the authors separate citation selection — being retrieved into the source list — from citation absorption, where a cited page actually contributes language, evidence or structure to the answer. A page can be selected constantly and absorbed rarely. Higher-absorption pages were longer, more structured, more semantically aligned with the query and richer in extractable evidence such as definitions, statistics and comparison tables. A gap matrix that only counts appearances in source lists measures the cheaper half of the problem.
Table 4: the source list
The fourth table lists every URL the engines cited in answers where the brand appeared, grouped by domain with a count. It is the least glamorous section and often the most actionable, because it names the specific third-party surfaces doing the work — a directory profile, a review page, a comparison article, an industry publication. Over several months it becomes a map of the brand's real source graph and the evidence base for deciding where next month's off-site effort should go.
What the report should not contain
Three things routinely appear in monthly AEO reports and should not. A single blended visibility score with no published formula: if the reader cannot reconstruct it, it cannot be audited. A website traffic graph presented as proof of AI visibility: AI-driven visits largely arrive without a readable referrer — one analysis of 446,405 visits found 70.6% landing in the 'Direct' bucket — so traffic is a poor, lagging proxy for a channel measured properly by citation, mention and recommendation rate. And week-over-week movement on a stochastic system: below about ten runs per prompt, a small change is noise, and reporting it as progress is reporting noise.