ARO Index
ARO Index Free Audit Book a call Pricing For Agencies Results Compare Insights Agency Console Affiliate Console
ARO Index ARO Index.
← Back to Insights
Article 4 · June 2026

Cited Is Not Recommended: The Metric AI Search Is Missing

The gap between being cited and being chosen is where most of the confusion in this space lives.
By Therese Grittner for the ARO Index · June 21, 2026

Most tools measuring AI search today count one thing: did the model mention you. Citation. It is the easy number to grab, so it is the number everyone grabs. But citation answers the wrong question. When a person asks an AI assistant "who should I hire for X near me," the model does not hand back a bibliography. It picks. It names a business, sometimes two. That act of picking is selection, and it is a different event from being cited.

The gap between the two is where most of the confusion in this space lives. A business can be cited constantly and selected rarely. It can be left out of the sources entirely and still be the one the model recommends. Cited and recommended are not the same thing, and measuring the first while claiming to understand the second is how a lot of brands end up confidently wrong.

Why selection is the number that matters

Citation tells you the model knows you exist. Selection tells you the model chose you over the alternatives when it actually had to answer a buyer. Only one of those moves revenue.

If you are a small business, you do not need to be in the footnotes of an AI answer. You need to be the answer. That is the thing worth tracking, and almost nobody is tracking it.

Local is also where the measurement gets hard, which is exactly why it forces a real method. A national brand gets asked about constantly, so it shows up across thousands of queries and even a crude measure catches it. A single-location business might appear in a handful of relevant queries total. At that scale, the model's natural run-to-run noise can swamp the signal entirely. You cannot get away with a loose measure locally. The small-business case is the stress test that demands the rigor described below.

Why one screenshot is anecdote, not data

Here is the part that trips people up. You ask the model your question, it recommends you, you screenshot it, you feel great. You ask again next week and you are gone. Neither result is the truth. Both are samples.

AI model outputs are non-stationary and noisy. The same prompt run multiple times does not return the same answer, and the answer shifts based on how the question is phrased, which model is asked, and when. In our own testing, a meaningful share of queries flip the recommended business based on phrasing alone, with no change to the business itself. One screenshot captures a single roll of that die. Treating it as a fixed score is the core mistake.

How to measure selection honestly

You cannot read a noisy, shifting signal as a single number. You can still measure it. You just have to measure it like a tracking poll instead of a thermometer.

The method comes down to three moves:

Run it as a rate, not a result.

Run the same query enough times to get an appearance rate: the share of runs where the business is selected. Put a band around it. The band is not a flaw in the measurement. The variance is the measurement.

Only trust moves that clear the noise floor.

If your rate wobbles inside the band, nothing happened. Act only on changes large enough to clear the noise. This is what separates a real shift from the model simply being the model.

Keep models separate. Never pool them.

A business can sit stable on one model and near-zero on another. Averaging every model into one blended score erases exactly the signal that matters most. Track a rate per model, then look at how much the models agree. Agreement across models is a stronger, more durable signal than a high score on any single one.

Why the method is the whole point

Run the queries cold and separate. Bundle several questions into one conversation and the model anchors on its own earlier answers, which quietly suppresses the variance you are trying to measure. The honest version is uncomfortable and a little expensive: separate, repeated, single-turn calls, per model, scored as a distribution.

That discomfort is the point. The brands treating one good screenshot as proof are not the careful ones. The careful ones are measuring the distribution and only moving on what clears the floor.

Where this goes

Selection is measurable. It is just harder than citation, which is most of why the industry defaulted to citation. As AI becomes the layer people use to find local businesses, the question stops being "did the model mention me" and becomes "how often does it pick me, on which models, and is that holding." That is a research question, and it deserves a research method rather than a screenshot.

Part of the Insights series: Article 1, Article 2, and Article 3 cover the tool landscape, score variability, and how AI recommendation measurement is evolving. The ARO Index publishes live AI recommendation rankings for local businesses. Full method on the methodology page.

Ask the Index×
Hi! Ask me about the ARO Score, how AI recommendation works, or what separates top-ranked businesses.