Your AI visibility score is lying to you
(Just not on purpose)
by Julian Rogers
Somewhere in the last year, "AI visibility score" became a line item in marketing reports. A number shows up in a dashboard, a client asks what it means, and everyone nods like it's as solid as a Google Analytics session count.
It's not. And if you're building strategy around these numbers as if they were, you're building on sand.
Let's be clear about why (because the reason most people assume is not the real reason).
It's not the tools. It's the machine.
The easy explanation is that AEO tools are new, the math behind them is a black box, and nobody's agreed on a standard yet. That's true, and it's worth flagging. Plenty of platforms won't tell you exactly how they weight a citation or score a "mention." Calculation transparency is a real, legitimate gripe.
But that's not the main problem. The main problem is the thing being measured.
AI-generated responses are not stable. Ask the same model the same question twice — same day, same phrasing, no changes on your end — and you can get two different answers, two different sets of cited brands, two different summaries of "the best option." The model isn't retrieving a fixed webpage and reporting back. It's generating a response, informed by probability, context and a training process that has some amount of randomness baked in by design.
That means an "AI visibility score" isn't measuring a fixed thing that changed. It's sampling an inherently variable process and reporting back a snapshot. Do that on a Tuesday and a Thursday and you'll get different numbers — not because your content changed, and not because the tool is broken, but because the underlying system doesn't produce the same output twice on command.
Treat these scores like a weather forecast, not a thermometer reading. Useful for direction. Dangerous if you treat the number itself as gospel.
Which AEO metrics to trust
If the exact score is soft, what's actually worth watching? A few things hold up better than others:
Trend over time, not a single snapshot.
Is your brand showing up more often across dozens of prompts run over weeks, not one query run once? Direction matters more than any single data point.Citation presence on your owned assets.
Are AI answers actually pulling from your site, your published content, your Google Business Profile, versus a competitor's or a third-party aggregator's? That's a more concrete signal than a composite "score."Category and query coverage.
Are you showing up across the range of questions your buyers actually ask, or only for one lucky phrasing? A single strong result tells you almost nothing about your real footprint.Consistency across repeated testing.
If you ask a version of the same question ten times and you show up seven of those ten, that's a meaningfully different position than showing up once. That ratio, tracked over time, is more honest than any single-run "visibility percentage."
None of these replace judgment. All of them are more trustworthy than a single number pulled from a single query run once.
Why manual prompt testing still matters
Automated tools are useful for scale, but the primary reason manual prompt testing earns a place in your AEO measurement process isn't speed or convenience — it's that automated tools sample a narrow slice of how real people actually phrase real questions. A human running your own set of prompts, in the actual language your buyers use, catches variation and context that a standardized tool query misses. It's the difference between checking your reflection in one mirror and walking around the room. You need both, but if you only have time for one, manual testing is what surfaces how you're actually showing up — not just how a tool's default query shows you showing up.
The foundation for AEO visibility
Here's the part that trips up a lot of practices chasing AI visibility: none of it works without the boring stuff done first.
HubSpot Academy's framework for AI visibility lays this out as a pyramid, and the base of that pyramid isn't AEO tactics, AI-specific content, or clever prompt engineering. It's a strong traditional search foundation. Solid technical SEO. Real authority signals. Content that actually answers the question it claims to answer.
You need these three aspects:
structured,
crawlable,
trustworthy.
These remain the fundamentals that have mattered since before anyone said "large language model" out loud.
AI systems are, at their core, trained on and informed by the web as it exists. This includes the same signals that have always separated credible sources from noise. Skip the foundation and chase the AEO layer on top, and you're decorating a house with no frame under it. It might look fine in a screenshot. It won't hold.
If your traditional SEO is shaky, fix that first. The AI visibility work you do after that will actually have something to stand on.
Metrics, prompts, pyramids — none of it matters much without a strategy tying it together. If you want the fuller picture on how SEO, AEO, and GEO actually fit together for a growing practice, we laid it out here.

