What your AI visibility score isn’t telling you

It’s difficult to have a conversation with any PR practitioner today without AI, GEO, AEO, and LLMs coming up.

When it comes to properly measuring AI visibility, it’s easy to get frustrated and simply start singing Old MacDonald’s E-I-E-I-O instead.

We have spent years building measurement programs on metrics that behave predictably. Run the same query against the same media database twice and you get the same result twice.

That reliability is a big part of why anyone trusts things like share of voice or sentiment enough to put it in front of a board.

Large language models don’t behave the same way. You and I can type identical prompts into the same model and get different answers. Anything built on top of that is a different kind of number, and treating it like just another media metric can be problematic.

I recently sat down with Khali Sakkas, CARMA’s Chief Insights Officer, who attended the AMEC Global Summit in Dublin where AI visibility was unsurprisingly a hot topic.

She is a useful guide since she has been active in establishing industry standards, including serving as a member of the working group behind the latest update to the Barcelona Principles. In addition, she leads on what CARMA actually puts in front of clients.

The hard part isn’t the technology

Khali is blunt about which piece of this is difficult, and it is not the plumbing.

«Obviously the connections to the LLMs through APIs and all those things is the easy part,» she told me. «The technical part is the easiest part of this.»

The real hurdle is that the research does not replicate. «You put in some searches, the same prompts into your ChatGPT or different LLMs, I put the same prompts in, we will get different answers,» she said. «That’s the fundamental challenge of the whole industry.»

The problem is compounded by the fact that nobody sees what people are actually asking. «We don’t know what people are putting into LLMs. No one does. It’s not publicly available.»

Search queries, autocomplete suggestions, and the questions people ask on social platforms can inform a reasonable guess about how real people phrase things. But it is still a guess, and the quality of any GEO measurement rests on how good that guess is.

«We’ve always been talking about how important it is to be able to repeat research for it to be valid,» Khali said. «This is not a repeatable process.»

Volume is part of the answer

If one prompt tells you almost nothing, the response is to run a great many of them.

«We need to make sure that we have the most variations of those prompts. We need to make sure that we’re prompting those LLMs many, many times,» she said. «And when I’m talking about many times, not 50 or 100 times, we’re talking about hundreds of thousands of times to make sure that we understand the patterns and the direction that it’s going.»

That gives you a practical test as a buyer. Ask anyone selling you an AI visibility number how many prompts sit behind it, how those prompts were chosen, and how often the set gets refreshed.

New models arrive constantly, and they do not all behave the same way. When someone asks how you are showing up in ChatGPT, the honest follow-up is which version of ChatGPT, running when.

Threading that needle, holding a method steady enough to compare results over time while staying responsive to a platform layer that keeps moving, is the real work. I do not see the pace of change on the AI provider side slowing down, so this is a challenge we will likely need to solve more than once.

Directional evidence, not hard numbers

Fortunately, AMEC has stepped in to help provide guidance. The industry group launched its GEO Principles at the Dublin summit in May, along with a practitioner’s guide. The framework looks at AI-led discovery across three connected areas:

  • the upstream reputation signals that feed the models,
  • whether your own content is structured and credible enough to be interpreted, and
  • what the models then actually say about you.

It also sets baseline evidence requirements, including repeatable prompts, documented methods, stated assumptions, and clear limitations.

The Principles treat AI outputs as directional evidence rather than absolute truth.

Khali put it plainly. «The results should be taken as a directional approach rather than quantitative, hard statistics,» she said. «We’re used to having metrics that are repeatable, that you get the same results every time, whether that be marketing analytics or earned media measurement. We’re used to that being quite solid. Here, we need to think about it more as a directional approach of what’s getting cited, what’s the tonality of those LLMs, what are the narratives, rather than hard and fast numbers.»

That is a harder story to tell internally than a single figure on a dashboard. But it is also the only version that survives the question of how the number was produced. An honest finding presented directly is worth more than a confident one that falls apart under even gentle examination.

The old rules didn’t go away

Despite the challenges of accurately assessing GEO visibility, the news for the PR industry in this era of AI is encouraging. Khali points to studies suggesting that somewhere between 60 and 80 percent of LLM responses are driven by earned media citations, which puts communications work at the center of a discovery process that used to belong to the SEO team.

«The champagne is still flowing,» she said, on the industry celebrating this resurgence. «The hype is real.»

None of that exempts anyone from proper methodology, and the fundamentals that governed measurement before still govern it now. The most recent update to the Barcelona Principles reworked the seventh principle around ethics, governance, and transparency in data, methodology, and technology.

Khali noted the first question that surfaces when a team tries to bring AI into its workflow is whether its data is accurate enough to support it.

She is also realistic about where most organizations are starting from. The updated principles came with supporting materials on purpose, because teams are at wildly different stages.

«It doesn’t matter where you start, we just have to start,» she said. «You don’t have to overwhelm your team or overwhelm your clients with confusing metrics.»

What to do with this

Take the direction seriously and treat the decimal points skeptically. Ask how the number was built and expect the methods to evolve alongside the underlying technology.

«As soon as we conquer this challenge, there’ll be more challenges coming along,» Khali said. «The ground’s moving, and we have to keep up with that.»

The organizations that handle this well over the next few years will not be the ones with the highest AI visibility score. They will be the ones that can explain where their number came from and understand what it actually means.

Speak with one of our experienced consultants about your media monitoring and communications evaluation today.