ARO Index
ARO Index Free Audit Book a call Pricing For Agencies Results Compare Insights Agency Console Affiliate Console
ARO Index ARO Index.
← Back to Insights
Article 5 · July 22, 2026

Why We Measure Cold: The ARO Index Methodology, 2026

Consumer AI search results change based on who is logged in. A locked, cold measurement is the only kind that holds still long enough to compare.
By Therese Grittner for the ARO Index · July 22, 2026

There is a growing body of evidence that ChatGPT Search is not one product. It is many products, served to different people at different times.

Technical SEO researchers documented this in July 2026. Jerome Salomon ran identical queries on a paid account and a free account from the same location in France and got different grounding pipelines. His paid account pulled results labeled from Bing. His free account triggered a live product carousel experiment. Suganthan Mohanadasan's discovery of the hidden "result_source" field confirmed the mechanism: the values returned were tied to the specific account running the test.

The grounding pipeline varies by user tier, model selected, reasoning effort, location, and account memory. OpenAI runs experiments on the search experience the same way Google runs experiments on its results pages. Two users asking the same question can get different answers for reasons that have nothing to do with the language model itself.

This matters for anyone trying to measure AI recommendations.

The measurement problem

If you test AI recommendations through a consumer interface, your results are shaped by your account. Your tier, your history, your location, and whatever experiment you happen to be enrolled in that day. You cannot separate the model's preference from your session's fingerprint. As Salomon put it, never assume your test is what most users will see.

A measurement that changes depending on who runs it is not a measurement. It is an anecdote.

How ARO Index measures

ARO Index audits use independent cold API calls. Every audit runs 3 buyer-intent queries across 4 models: ChatGPT, Claude, Gemini, and Perplexity. That is 12 reads per audit. Each call is a fresh context with no account, no memory, no history, and no interface-layer experiment. A business only counts as selected by a model when it appears in 2 of 3 queries for that model.

Cold calls are the control condition. They isolate the one thing that can be measured consistently: what the model itself selects when nothing else is influencing the answer. The method is reproducible. Anyone with API access and our locked query bank can run the same test and check the numbers.

What this method does not capture

Honesty about limitations is part of the method. A cold API call is not identical to a logged-in consumer session. Interface layers add search grounding, personalization, and live experiments on top of the base model. A business selected in our audits may surface differently for a specific user, and the reverse is also true.

So ARO Index does not claim to predict any single user's screen. It measures the model's baseline selection behavior. Think of it as the signal underneath the noise. The consumer interface adds variables no one can hold constant. The cold call removes them.

The three questions every buyer should ask

A July 2026 industry debate between Evertune and Profound put three questions on the table that most measurement vendors avoid. Here are our answers.

Who picked the prompts, and why those?

ARO Index runs a locked query bank of cold buyer-intent queries, built per category and published with each report. The same queries run for every business in a market. No business, client or not, gets custom prompts. Locked queries are what make scores comparable across a city.

How do you handle noise?

Single AI responses are volatile. That is why no single response counts. Each audit runs 3 queries across 4 models, 12 independent reads, and a model only counts as selecting a business when it appears in 2 of 3 queries for that model. We also measure binary selection, was the business chosen or not, rather than share-of-voice percentages. A yes-or-no signal with a majority vote carries a lower noise floor than a percentage built from pooled mentions.

How do you separate content effect from model drift?

We do not claim to. Every ARO audit is timestamped, and published figures are locked to dated snapshots. When a score changes between audits, the honest answer is that model behavior and site content can both move it. Anyone claiming to cleanly separate the two is selling precision this category has not earned.

Disclosure

ARO Index is the research arm. Its founder also operates TaG Makes, an implementation service. That is a conflict worth naming, and here is how it is handled: leaderboard rankings come from the same locked methodology for every business, client or not. The audit does not know who pays us. Scores cannot be purchased, and placement cannot be bought.

Why this is the standard

The July 2026 findings do not weaken interface-based tracking tools. They end the argument. If the test environment can silently change what you observe, interface testing cannot produce comparable data across businesses, cities, or time. Controlled, cold, repeated measurement can.

That is why every number ARO Index publishes comes from the same locked methodology. The audit does not create your score. It reveals it.

Sources: Jerome Salomon (Technical SEO, AI Search), LinkedIn, July 2026. Suganthan Mohanadasan (AI SEO Research, Snippet Digital), LinkedIn, July 2026.

Part of the Insights series: Article 1, Article 2, Article 3, and Article 4 cover the tool landscape, score variability, how AI recommendation measurement is evolving, and the difference between citation and selection. The ARO Index publishes live AI recommendation rankings for local businesses. Full method on the methodology page.

Ask the Index×
Hi! Ask me about the ARO Score, how AI recommendation works, or what separates top-ranked businesses.