The field that measures how AI talks about businesses is moving faster than most of the businesses it describes. Six months ago, asking an AI model one question per platform told you something real: did it recommend this business, or didn't it. That answer still matters. But the models have changed how they read a business, the questions buyers ask have multiplied, and the method has had to grow up to keep pace. What was current in January is already coarse by June.
AEO (answer engine optimization) and GEO (generative engine optimization) describe optimizing a business to appear in AI-generated answers. AI recommendation research measures something downstream: when an AI is asked to recommend, which business does it actually choose. Appearing is visibility. Being chosen is selection. The ARO Index measures selection. That distinction is the whole reason the method changed.
A single query is a single camera angle. It tells you what one phrasing surfaced on one day, and it hides how much the answer depends on how the question was asked. As the models grew more sensitive to phrasing, one angle stopped being honest. So every business is now read through three queries per model, not one. Not to track prompts, and not to inflate a number. Three queries show how a business is interpreted across the way real people actually ask, inside the industry it competes in. The score reflects selection across that range, not a single lucky phrasing.
Here is what that surfaced, in the open. Sailor Craft Knots, a one-person maker in Charleston, scores 76 on "best british slip leash in Charleston SC" and 95 on "custom dog leashes in Charleston SC." Same business, same week, two questions, nineteen points apart. One query would have reported one of those numbers and called it the truth. Three reports the real shape: where this business is already chosen, and where the opportunity still sits.
Tech Week sharpened the question. Sitting with people at the top of their fields, and running a round of follow-up research afterward, pushed a question we'd been circling into the open: are we measuring whether AI mentions a business, or whether AI chooses it. Those are not the same thing, and the difference is the entire category. The community made it sharper still. The questions people brought to the audit tool, the pushback on early scores, the "why did this change between runs" conversations, all pointed at the same gap and helped close it. Research gets tested in public. This one is.
What is the difference between AEO and AI recommendation measurement?
AEO measures whether a business appears in AI answers. AI recommendation measurement asks a different question: when the AI is told to recommend one option, which one does it pick. Appearance is not selection.
Why measure three queries instead of one?
Because AI answers shift with phrasing. One query measures one phrasing on one day. Three queries measure how a business is interpreted across the range of ways buyers actually ask, which is closer to real-world behavior.
Does an AI recommendation score change over time?
Yes. The models update, the competition changes, and the questions evolve. A score is a reading of a moment, not a permanent grade. That is why the index is dated and re-run, not set once.
Part of the Insights series: Article 1 and Article 2 cover the tool landscape and score variability. The ARO Index publishes live AI recommendation rankings for local businesses. Full method on the methodology page.