Answer Engine Optimization is the practice of improving the probability that AI engines mention, cite and recommend a brand. The load-bearing word is probability. AI answers are synthesized fresh from many sources, differ between engines, and shift with every model update, so no provider can promise a fixed placement or a set number of AI mentions. Almost every AEO agency red flag is a variation on one root error: treating a probabilistic, third-party system as if it were a deterministic dial the agency can turn. Once you see that, the tells become easy to name.
The market's own referees are blunt about it. In August 2025, Google's John Mueller warned that "the higher the urgency, and the stronger the push of new acronyms, the more likely they're just making spam and scamming." That is a useful filter: urgency plus jargon, minus evidence, is the shape of a pitch designed to close before you can check it. Two red flags do the most damage, so they get their own sections below; a fuller catalog follows.
Red flag one: the guarantee trap
The loudest tell is a guaranteed outcome — "we guarantee you'll appear in ChatGPT," "guaranteed AI citations," "a set number of AI Overview appearances." As one 2026 agency-selection guide puts it plainly, a red flag is when "they promise guaranteed rankings, citations, or mentions," because a credible agency can only work on "increasing your chances of inclusion." No agency controls how a model selects sources, so a guarantee is either a misunderstanding of the channel or a setup to redefine success later. The dedicated companion piece on why no honest AEO agency guarantees a placement walks through what AEO can and cannot move; for this catalog, the rule is simple: treat certainty as the red flag, never the selling point.
Red flag two: the black-box dashboard
The subtler and now more common tell is a polished dashboard that reports a single AI-visibility score with no way to inspect it. It says you are "number four in your category," "up two spots this week," or "17% visible" versus a competitor's 31%. The problem, as one engineering analysis of these tools argues, is that "the precision is made up." AI engines are noisy, personalized, geographic and nondeterministic — even a temperature-zero model call is "not perfectly stable in production" — so any single answer is "one sample from a distribution," not a ranking. Without the prompt list, runs per prompt, geography, model, account state and scoring formula, the dashboard is showing a constructed metric. Over forty tools now sell 'AI visibility,' and most "track things that don't correlate with traffic or conversions." A black-box dashboard produces the sensation of control without the evidence to act on — the exact opposite of measurement.
The fix a buyer can demand is transparency, not a prettier chart. A credible report separates brand mentions from citation sources, shows the prompt and the answer behind a number, and reports citation accuracy, not just citation volume. If a team cannot drill from a headline score into the individual prompt, answer, source and timestamp, treat the dashboard as a black box, not a partner. What a good AI-visibility report contains, section by section, sets the acceptance bar in detail.