ARO Index
ARO Index Free Audit Book a call Pricing For Agencies Compare Insights Agency Console Affiliate Console
ARO Index TaG Makes
ARO INDEX METHODOLOGY

The ARO Score Method (v3 · July 2026)

What the ARO Score measures, how audits are structured, and why the method changed in 2026.

Research by Therese Grittner ARO Index Charleston, SC
THE ARO SCORE METHOD · V3 · JULY 2026
What we measure

Selection, not mentions.

When an AI is asked to recommend a business in a category and city, does it choose this one. Being named is visibility. Being chosen is selection.

Two call types, one clear line

Evaluation calls and cold-selection calls answer different questions.

Evaluation calls (the ARO Score below) include the business's own site content by design and ask a model to judge whether it would select that business for a real buyer query. Because the model is working from real site content instead of guessing, every audited business starts from a score floor rather than a blank page. This measures whether a business is recommendable.

Cold-selection calls (used in ARO Index research reports) fire the exact same locked queries with no site content supplied and no business name planted in the prompt, and only count businesses the model volunteers entirely on its own. This measures whether a business is actually recommended. See the Observed-Selection Census method below for the full breakdown.

How an audit runs

12 independent reads per audit.

Each audit runs 3 separate queries per model across 4 models (ChatGPT, Claude, Gemini, Perplexity), each including the business's own site content and live web retrieval, producing 12 independent reads. Queries are never bundled, so models cannot self-anchor and manufacture false agreement.

Model Access Surface

What each model is actually working from.

Model API model string Retrieval state Confirmed date
ChatGPT gpt-5.4 Live web retrieval (web_search tool) 31 Jul 2026
Claude claude-sonnet-4-6 Live web retrieval (web_search tool) 31 Jul 2026
Gemini gemini-3.1-pro-preview Live web retrieval (google_search tool) 31 Jul 2026
Perplexity sonar-pro Search-grounded at provider level 31 Jul 2026

Confirmed by direct audit of the audit runner source code.

All four models run with live web search enabled on every audit call, using the business's own site content alongside what each model retrieves on its own. The ARO Score reflects that combined read, not training data alone.

v3 grounded became the default method for every new audit on July 31, 2026. Audits completed before that date remain on v2 and are labeled accordingly. v2 and v3 scores are not directly comparable, since v3 adds live web retrieval that v2 did not have. Businesses move to v3 automatically as their next audit completes.

Gemini's reasoning effort is fixed to a low, explicit setting on every audit call. Left unset, the model defaults to its highest reasoning tier, which increases latency and output token volume without changing the scoring rubric it is asked to apply.

How the score is built

Majority across queries, not a single read.

Per model, a 2-of-3 majority across the three queries decides whether that model recommends. The score reflects selection across the full range of queries, not a single lucky one.

Why three queries

One phrasing is one camera angle.

Models are sensitive to how a question is asked. Three separate queries show how a business is interpreted across how buyers actually search.

Query phrasing

Three intent angles, not three synonyms.

Each locked query bank carries 3 phrasings built on the same three buyer-intent angles: a direct query ("best [category] in [city]"), a referral query, and a seeking query ("looking for a [category] in [city]"). This is the same 3-phrasing structure used for both evaluation calls and cold-selection calls, so the two tracks above are always asking about the same underlying buyer intents.

The referral phrasing varies by category type. Place categories (restaurants, hotels, tours, coffee shops, and similar) use "can you recommend a [category] in [city]." Provider categories (professionals and service providers people hire directly) use "who would you recommend for [category] in [city]." Both ask the same thing in the way a real buyer would actually phrase it for that kind of business.

Why we run multiple queries

Not all models respond the same way twice.

A June 2026 variance study across five markets and four models found Gemini and Perplexity stable enough for a single read -- their top picks held consistent across repeated queries on the same business. ChatGPT and Claude were less stable, especially in crowded categories where many businesses have similar web presence.

The response: 3 separate queries per model, with a 2-of-3 majority required for a recommendation to count. One lucky answer does not move your score.

ARO Score™ formula

Four factors, published weights.

ARO Score™ is the diagnostic composite inside the ARO Report. It is 0 to 100 and combines four factors:

  • Recommendation rate (40%)
  • Position strength (20%)
  • Confidence (20%)
  • Signal quality (20%)

Recommendation rate carries the most weight because being chosen is the point. The other three factors reflect how prominently, how consistently, and how reliably that selection holds. ARO Score™ explains what may be driving a business's ARO Index Score. It is not the public benchmark itself.

ARO Index Score formula

The public benchmark, rebuilt on observed outcomes.

The ARO Index Score is the public benchmark used to rank and compare businesses within a market and category. As of July 2026, it is calculated entirely from observational inputs, not from the ARO Score™ diagnostic composite. Individual reads come from the audit method current at the time of each audit (v2 before July 31, 2026, v3 grounded after). The ranking formula below changed in July 2026 and is applied to both.

The ARO Index Score is 0 to 100 and combines four factors:

  • Recommendation rate (45%)
  • Cross-model consistency (20%)
  • Recommendation position (25%)
  • Recency (10%)

Recommendation rate carries the most weight because being recommended is the point. Cross-model consistency and recommendation position reflect how reliably and how prominently that recommendation holds across ChatGPT, Claude, Gemini, and Perplexity. Recency weights recent audits more heavily than older ones.

Position strength

Where you appear in the list matters.

A recommended business named first earns more weight than one named fifth. Higher ranks earn descending credit down the list. A business that is recommended but given no clear rank receives a neutral middle weight. This rule is fixed and applied identically to every business.

Signal quality

Four grouped factors, not a checklist.

Signal quality captures the site-side factors that affect how well AI platforms can read and trust a business. Four groups:

  • Content clarity
  • Technical crawlability
  • Structured data
  • Authority signals

No individual dimension score is published. The aggregate drives the component weight.

Cadence

Monthly by market.

Each market is re-audited monthly. Rankings update when new audits complete.

Conflict-of-interest rule

No business can pay to change its score.

Weights are fixed and applied identically to every business, client or not. Buying a subscription changes what advice you receive; it does not touch the audit engine or change how your score is calculated. Measurement and services are kept separate by design.

Data currency

A score reads a moment.

Scores are dated and re-run. Models update, competition shifts, questions evolve. A score reads a moment; it is not a permanent grade.

The original single-query method (v1, through early 2026) is preserved below for reference.

OBSERVED-SELECTION CENSUS METHOD · RESEARCH REPORTS
Used in published research (e.g. Charleston 2026)

A separate, stricter measurement than the ARO Score.

The ARO Score above audits a business against its own category queries. The Observed-Selection Census, used in ARO Index research reports, asks a different question: when a real buyer types an unprompted query with no business name planted in it, which businesses does the model name on its own. Every published research number (Charleston 2026 and future city reports) traces back to this method, not the ARO Score formula above.

Two lenses, not one number

Direct audit vs. cold market query.

Direct audit: run each business's own category queries across the four models and record whether it gets named. In the Charleston census, 76.2% of the 688 audited businesses surface at least once this way.

Cold market query: run the unprompted "best [category] in [city]" a buyer would actually type, with no business name planted in the prompt, and record which businesses the models name on their own. In the Charleston census, only 12.9% are ever selected this way.

A business can pass the direct-audit lens and still fail the cold-query lens. That gap is the finding: recommendable is not recommended. Reports lead with the cold-query number because it is the one that reflects what an unprompted buyer actually sees.

Retrieval state for the Charleston 2026 census: cold market queries ran without live web retrieval on ChatGPT, Claude, and Gemini. Perplexity was search-grounded at the provider level. A business not named in this census was not named without live retrieval. A grounded (v3) rerun of the census is planned. When it runs, the parametric-to-grounded delta will be published separately.

What counts as selected

Matched by normalization function, not eyeballed.

A business counts as selected when a model names it unprompted in a cold market query and that name (or domain) matches the audited business, using the same normalization functions that built the entity corpus for that market. The exact functions and the full matching query are published below, so anyone can rerun it.

Matching runs on normalized name or normalized domain, never domain alone — a domain-only rule would undercount, since not every cold-query mention carries a domain.

The exact matching query · normalization functions · reruns live
Normalization functions (live in the database)
CREATE OR REPLACE FUNCTION public.phase2a_normalize_name(raw text)
 RETURNS text LANGUAGE sql IMMUTABLE AS $function$
  SELECT trim(
    regexp_replace(
      regexp_replace(
        regexp_replace(
          regexp_replace(
            regexp_replace(
              lower(trim(raw)),
              '^the\s+', ''
            ),
            '''s(\s|$)', '\1', 'g'
          ),
          '[''.,]', '', 'g'
        ),
        'barbecue', 'bbq', 'g'
      ),
      '\s+(oyster bar|real estate|restaurant|llc|inc)\s*$', ''
    )
  );
$function$;

CREATE OR REPLACE FUNCTION public.phase2a_normalize_name_aliased(raw text)
 RETURNS text LANGUAGE sql IMMUTABLE AS $function$
  SELECT CASE
    WHEN lower(trim(raw)) IN ('hall''s chophouse', 'halls chophouse') THEN 'halls chophouse'
    ELSE phase2a_normalize_name(raw)
  END;
$function$;

CREATE OR REPLACE FUNCTION public.phase2a_normalize_domain(raw text)
 RETURNS text LANGUAGE sql IMMUTABLE AS $function$
  SELECT NULLIF(
    regexp_replace(
      regexp_replace(lower(trim(raw)), '^https?://', ''),
      '^www\.', ''
    ),
    ''
  );
$function$;
Matching query (Charleston 2026 example, reruns live)

Full runnable version, including the audited-cohort definition and one documented manual alias, is published in _config/METHODOLOGY.md. Core matching logic:

select count(distinct ac.id) as selected_count
from audited_businesses ac
join selected_entities se
  on phase2a_normalize_name_aliased(ac.business_name)
     = phase2a_normalize_name_aliased(se.canonical_name)
  or phase2a_normalize_domain(ac.domain)
     = phase2a_normalize_domain(se.canonical_domain);

Where selected_entities is the deduplicated canonical-entity layer built from clean cold-query mentions (excludes off-category bleed and any retro-mined pilot data), and audited_businesses is every business with a current, public audit in the market.

Locked example: Charleston 2026

688 businesses audited · 89 ever selected in a cold buyer query · 12.9% selection rate · 1,616 clean cold mentions · 1,047 distinct entities · 43 categories · data current July 8, 2026. Full report: /research/charleston-2026.

Used in longitudinal comparisons (e.g. Model Shift reports)

Research cohorts are frozen snapshots.

A research cohort locks a list of businesses on one date, so a later comparison measures the same businesses over time. Each cohort records a membership rule, a snapshot date, and a locked count. The frozen list is authoritative. The rule describes how the list was built. It does not reproduce the list on demand.

charleston_vol1: best-known rule — audited, score version v2, live audit environment, public, 3 or more models, most recent qualifying audit before the snapshot. Snapshot Aug 26, 2026. 687 domains. Locked. The live version of this query drifts as audits accrue — 688 on July 8, 685 on Sept 13. The frozen list defines the cohort, not the live query. Running the rule today does not return 687, and it isn't supposed to.

nashville_vol1: same rule, scoped to Nashville, cutoff before July 27, 2026 — the day the Nashville July census fired. Snapshot Sept 13, 2026. 481 domains. Locked.

A cohort built after its own census already ran is disclosed as retroactively constructed with a pre-census cutoff. It is never called pre-registered.

V1 archive · original single-query method · through early 2026
Section 1

About the ARO Index

The ARO Index is a live market dataset tracking which businesses AI platforms actually select when buyers ask for local service recommendations by city and category. It measures selection behavior, not mentions. A business can be mentioned without being selected. ARO Index records selection.

Section 2

What We Measured

ARO Index tracked AI recommendation selection across four platforms: ChatGPT, Claude, Gemini, and Perplexity. For each market and category, a standardized query was submitted to all four platforms. The responses were recorded, analyzed, and used to produce ARO Scores and market rankings.

Rankings reflected which businesses were selected most consistently and prominently across platforms within a defined audit period.

Section 3

How Audits Worked

Audits were conducted on a rolling cycle. Each audit covered a defined city and service category. A single query was submitted in a standardized format across all four AI platforms. Results were recorded at the time of submission.

ARO Index does not rely on scraped data, third-party aggregators, or predictive modeling. All data is collected through direct platform queries conducted by the ARO Index research team.

Section 4

Audit Scope

Approved markets:

  • Charleston, SC
  • Nashville, TN
  • Atlanta, GA
  • Denver, CO
  • Miami, FL
  • Sacramento, CA
  • Fresno, CA

Additional markets are added on a rolling basis. Each market is organized by service category.

Section 5

ARO Score

The ARO Score is a composite measurement scored 0 to 100. It reflected how consistently and prominently a business was selected by AI platforms when buyers asked for recommendations in its city and category. A higher score indicates more consistent selection across platforms and queries.

Scoring methodology is proprietary. The ARO Score is a trademark of TaG Makes.

Section 6

Data Currency and Versioning

ARO Index data reflects point-in-time audit results. AI platforms update their models continuously, which means recommendation behavior can shift between audit periods. Each report references its specific audit period and dataset.

Current rankings are always available at aroindex.com.

Section 7

Research Lead

ARO Index research is conducted by Therese Grittner, founder of TaG Makes, based in Charleston, SC. Therese Grittner created the ARO Index methodology and oversees all audits, scoring, and market reporting.

Section 8

Limitations

ARO Index measures AI platform behavior at the time of each audit. Results reflect the platforms' recommendation behavior during that specific period and may not represent behavior at other times.

Sample sizes vary by market and category. Posts with samples under 100 businesses are labeled as exploratory or directional.

ARO Index does not make causal claims. All findings are observational and correlative unless otherwise stated.

Section 9

Contact

For research inquiries: therese@aroindex.com

For current rankings and market data: aroindex.com

Ask the Index×
Hi! Ask me about the ARO Score, how AI recommendation works, or what separates top-ranked businesses.