ARO Index Research Report · Vol 3

ARO Index Vol 3: The September 2026 Model Shift

By Therese Grittner ARO Index Published September 23, 2026 Cold-query observed selection, 4 AI models Data current as of Sept 14, 2026

Headline

Between July and September 2026, the number of cohort businesses that AI models actually recommended rose in both markets. Charleston went from 89 of 687 to 130 of 687 on the same 43 categories. Nashville went from 87 of 481 to 152 of 481 on the same 82 categories.

That comparison is not a same-method time series. The July censuses ran without live web retrieval on ChatGPT, Claude, and Gemini. The September census ran with live retrieval on all four models. Section 1 is the parametric-to-grounded delta the ARO Index methodology page has promised since July, published here for the first time. Perplexity, the one model whose retrieval condition did not change, held at 77 of 687 in Charleston across both periods. The rise on the other models is retrieval plus time, and this report cannot separate the two.

Section 4 is the controlled test. Both of its panels ran with live retrieval on all four models, on the same days, with the same categories. When we re-ran the September census on the new default-tier models (the ones a consumer actually talks to), selection went slightly down, not up. Charleston 151 to 147. Nashville 170 to 162. ChatGPT was the only model that gained. Gemini drove the decline. Perplexity, run unchanged as the control, barely moved.

Decision: no methodology version bump. The live pipeline moved to the current consumer base models on Sept 16 as a dated panel change (see 4.6).

Suggested citation

ARO Index, ARO Index Vol 3: The September 2026 Model Shift, published September 2026.

Cite this finding

Between July and September 2026, the share of a locked 687-business Charleston cohort recommended by AI models rose from 89 of 687 (13.0%) to 130 of 687 (18.9%) on the same 43 categories, with the July baseline run without live web retrieval and the September census run with it. On a controlled same-day test of the new default-tier model panel, selection fell slightly: Charleston 151 to 147, Nashville 170 to 162. ARO Index, Vol 3: The September 2026 Model Shift, aroindex.com/research/model-shift-sep2026. Data current September 14, 2026.

Media contact

Therese Grittner, ARO Index, therese@tagmakessc.com

Section 1: July vs September, Same Categories, Same Cohort

These figures compare the July census (no live retrieval on ChatGPT, Claude, Gemini) to the September pre-shift census (live retrieval on all four models). Both periods use the same cohort and the same categories. Read the deltas as retrieval condition plus time, not as time alone. Section 3(a) has the detail. Gemini is excluded from Sections 1 and 2 only: its July responses averaged 86 to 93 characters (a known incident, the responses were hollow). September Gemini is healthy and is included from Section 4 onward.

1.1 Charleston (43 shared categories, 3 models)

July 7, 2026Sept 10, 2026
Cohort businesses selected89 of 687130 of 687
Selection rate13.0%18.9%

Retained (selected both periods): 64. New in September: 66. Dropped: 25. Arithmetic: 64 + 25 = 89. 64 + 66 = 130.

Per model, cohort businesses selected (distinct, name-or-domain match):

ModelJulySept
ChatGPT1951
Claude2184
Perplexity7777

Why 43 categories. The July 7 Charleston batch (mkt_pilot_charleston_2026_07) ran 43 categories. The September census ran the full frozen 82-category bank. Every one of the 43 July categories exists verbatim in the September bank, so the comparison is restricted to that shared set. An earlier uncontrolled comparison reported 89 to 200; that figure mixed a 43-category baseline against an 82-category run and is retired.

Coverage expansion, reported separately and never framed as improvement. 102 cohort businesses were selected in at least one of the 39 September-only categories. 70 of those were selected only there, never under a shared category in either period. Those 70 are the category bank asking more questions, not existing categories getting more generous. 130 + 70 = 200, which reconciles exactly to the retired uncontrolled total.

1.2 Nashville (82 categories, 3 models)

July 27-28, 2026Sept 10, 2026
Cohort businesses selected87 of 481152 of 481
Selection rate18.1%31.6%

Retained: 73. New in September: 79. Dropped: 14. Arithmetic: 73 + 14 = 87. 73 + 79 = 152. This July baseline is batch census_jul2026_nashville_v2 (82 categories, July 27-28). It is not the 22-category, 264-query Nashville first look (39 of 485) reported in Vol 2, which came from a different batch; the two figures are not comparable.

Per model:

ModelJulySept
ChatGPT2268
Claude24102
Perplexity6879

Both Nashville batches cover the identical 82-category bank, so no category restriction was needed.

Section 2: Businesses That Dropped Out

Selected in July, not selected in September, same categories both periods.

2.1 Charleston (25)

Listed with the July category each was selected under. Two businesses were selected under more than one category in July.

2.2 Nashville (14)

Section 3: Caveats on Sections 1 and 2

Disclosed, not resolved.

(a) The retrieval condition changed between periods. This is the largest confound in Section 1. The July batches (mkt_pilot_charleston_2026_07, census_jul2026_nashville_v2) ran through the census pipeline's no-tools path on ChatGPT, Claude, and Gemini. The September batches (census_sep2026_*_pre) ran through the live-retrieval path on all four models. Perplexity retrieves by design and was grounded in both periods. Verified Sept 22, 2026 two ways: the July worker source (commits 47d45a88 and 7a92a389) contains no retrieval branch, and market_runs.usage_raw is null on 100% of July rows and populated on 98.8% to 100% of September rows, which only the retrieval branch writes. This disclosure was added Sept 22, 2026 after the draft was complete; the Sept 15 draft did not carry it.

The time gap compounds it. Charleston: July 7 to Sept 10 is 65 days. Nashville: July 27-28 to Sept 10 is about 6 weeks. Both periods ran on current, non-frozen model strings. The Section 1 deltas capture retrieval condition, time drift, and model drift together, and this report does not separate them. Section 4 is the controlled test: both of its panels ran with live retrieval.

(b) Model strings as stamped. The September baseline (census_sep2026_*_pre) ran ChatGPT gpt-5.4, Claude claude-sonnet-4-6, Gemini gemini-3.1-pro-preview, Perplexity sonar-pro. The July Nashville batch is stamped ChatGPT gpt-5.4-mini-2026-03-17, Claude claude-sonnet-4-6, Gemini gemini-3.5-flash, Perplexity sonar. The July 7 Charleston batch predates version stamping and is recorded as unverified. Only Claude's string is confirmed identical across periods. Readers should treat Section 1 as a same-generation comparison on whatever strings were live on each date, not a fixed-string comparison.

(c) Longer answers are a symptom of (a), not a separate cause. Claude's average response length grew from 1,183 to 3,436 characters between July and September. ChatGPT and Gemini responses grew 20x to 30x. More text means more candidates named per call. Mention volume in September was about 2.3x July, explained by response length, not extractor changes. Live retrieval puts fetched web content into the model's context, and longer, more specific answers are what that produces. Perplexity, whose retrieval condition did not change, did not grow the same way and held at 77 of 687 in Charleston. Some of the Claude and ChatGPT selection jump in Section 1 is retrieval, not better targeting by the models.

(d) Cohorts were constructed after the fact. Both cohorts follow the frozen-snapshot rule in methodology section 8. charleston_vol1 (687) is a best-known reconstruction: the documented query returns 688 on July 8 and 685 on a Sept 13 rerun, and neither reproduces 687. The locked list is authoritative; the rule is descriptive. nashville_vol1 (481) was built Sept 13 with a pre-census cutoff (audits created before July 27) and was never pre-registered.

(e) Nashville entity resolution ran out of order. The September Nashville side was entity-linked Sept 11. The July Nashville side was not linked until Sept 13, after the comparison had already been attempted once. Final linkage: 3,168 of 3,170 mentions (99.94%). Raw name matching without the entity layer was tested and rejected: it returns 167 versus 152 on the same September batch, a real 10% gap.

(f) Pre-June scores are not comparable. Anything scored under the retired V2 parametric methodology cannot be compared to this pipeline.

Section 4: The Controlled Model Test

4.1 Setup

The "before" panel is the September pre-shift census (Sept 10, current strings, all four models healthy). The "after" panel re-ran the identical categories, variants, and matching on new model strings on Sept 14. Same cohorts. Charleston restricted to the same 43 shared categories as Section 1. Nashville on the full 82.

After-panel strings (default tier, what a consumer actually uses): ChatGPT gpt-5.6-luna, Claude claude-sonnet-5, Gemini gemini-3.8-flash, Perplexity sonar-pro (unchanged, the control).

Gemini exception, disclosed: the Gemini leg ran gemini-3.8-flash, Google's current GA API model. On Sept 22, 2026, the free Gemini consumer app defaulted to 3.6 Flash on two separate accounts. Whether 3.6 Flash was the consumer default on Sept 14 is not confirmed. The Gemini figures in this section are default-tier API results, not verified consumer-default results.

Retrieval condition: both panels ran with live web retrieval enabled on all four models. The before panel and the after panel differ only in model strings. Verified Sept 22, 2026 from market_runs.usage_raw on every leg of census_sep2026_*_pre, model_test_free_20260914, and model_test_v4_20260914.

Run-date split, disclosed: Gemini and Perplexity legs landed by about 16:14 UTC Sept 14. Default-tier ChatGPT and Claude legs landed 21:27 to 22:55 UTC Sept 14. Row counts verified in market_runs: 375 of 375 per model per batch. Unresolved mentions after entity resolution: 0 on both markets.

4.2 Results, All Four Models

Charleston (43 shared categories): 151 to 147.

Retained 122. New 25. Dropped 29. Arithmetic: 122 + 29 = 151. 122 + 25 = 147.

ModelBefore (Sept 10)After (Sept 14)Change
ChatGPT5158+7
Claude84840
Gemini7972-7
Perplexity (control)7776-1

Nashville (82 categories): 170 to 162.

Retained 139. New 23. Dropped 31. Arithmetic: 139 + 31 = 170. 139 + 23 = 162.

ModelBefore (Sept 10)After (Sept 14)Change
ChatGPT6882+14
Claude102101-1
Gemini10089-11
Perplexity (control)7978-1

Reproduction check: before trusting the method for new numbers, the Gemini and Perplexity figures already on record were re-derived independently against live data. Exact match in both markets.

Reading it. Perplexity did not change strings and moved by one in each market. That is the noise floor. ChatGPT's gain (+7 Charleston, +14 Nashville) clears it. Gemini's loss (-7, -11) clears it in the other direction. Claude is flat. Net, the new generation recommends slightly fewer cohort businesses than the old one, and the control confirms that is not a method artifact.

4.3 Dropped in the Model Test

Selected on the Sept 10 strings, not selected on the Sept 14 strings, same categories.

Charleston (29):

Nashville (31):

Note: two distinct cohort entries both display as "Nashville SEO Agency" (hortongroup.com and astute.co). Verified as two different businesses sharing a scraped name, not a duplicate. Source: ClickUp comment 1000410000020126, re-derived live Sept 15, 2026.

4.4 Flagship Models, Tested and Set Aside

A first attempt at the "after" panel used flagship strings (ChatGPT gpt-6-astra, Claude claude-fable-5-1). That run was stopped at 84 of 375 ChatGPT rows and 87 of 375 Claude rows in Charleston, and 0 rows in Nashville, when API credits ran out. It was not resumed, on purpose: flagship calls cost roughly 7x the default tier, and flagship models are not what a consumer talks to in the free or default product. The default-tier run replaced it as the "after."

Where flagship and default rows both exist (Charleston, 84 category-by-variant combos), the paired result is mixed and directional only: ChatGPT default 45 versus flagship 34; Claude default 66 versus flagship 70. One model lower on flagship, one slightly higher. No conclusion is drawn from it.

Control integrity, verified Sept 16: every Perplexity row in the before and after panels (246 + 246 + 375) ran on the Chat Completions sonar-pro path before the Sept 14 Agent API migration. Zero rows resolved to any other vendor. Source: ClickUp comment 1000410000020213.

A related finding from the migration probe, disclosed here because it bears on how these panels are read: Perplexity's Agent API preset:"low" tier silently returned openai/gpt-5.6-luna as the answering model on 15 of 15 test calls, with HTTP 200 and normal-looking output. The ARO Index pipeline pins a fixed model string on every call and never uses preset. From wave 2, the Perplexity control string is perplexity/sonar (sonar-pro sunsets Sept 27, 2026), and that change will be disclosed in the next census.

One observation worth tracking, not yet a finding: in a single-call probe, Claude returned 26 candidate businesses where ChatGPT returned 7 and Gemini 15. If that ratio holds across batches, the pattern of Claude naming far more businesses than ChatGPT is structural to how the models answer open-ended recommendation queries, not version drift.

4.5 Cost

Default-tier after-panel, all four legs, real billed: ChatGPT $1.43, Claude $17.60, Gemini $1.15, Perplexity $4.26. Total $24.44.

Abandoned flagship spend, kept as a separate line and never blended: Claude $20.23 plus ChatGPT $14.76, $34.99, partial, never completed. ChatGPT flagship figure corrected Sept 22, 2026 from $14.86 to $14.76 against the OpenAI usage dashboard, Sept 14 UTC, grouped by line item (gpt-6-astra input $12.06, output $2.22, cached input $0.36, cache writes $0.12).

OpenAI web search tool calls, disclosed separately: $7.93 for 793 searches across 463 requests on Sept 14 UTC. The dashboard bills search calls as one line for the whole day and does not split them by model, so this cost spans both the default-tier and flagship ChatGPT legs and is not attributable to either. It is not included in the ChatGPT token figures above. Total OpenAI spend for Sept 14 UTC, all line items: $24.21.

4.6 Decision

NO-GO on a model version bump. Both markets show a small net decline on the new default tier (151 to 147, 170 to 162). ChatGPT is the only model with a real gain. Gemini drives the decline. Perplexity as the unchanged control confirms the decline is not a method artifact. Flagship data, where it exists, does not support a bump either. As a research finding, this report does not support moving the census to the new strings.

Production note, disclosed: on Sept 16, 2026, the live ARO Index audit pipeline moved to the current consumer base models (gpt-5.6-luna, claude-sonnet-5, gemini-3.8-flash, perplexity/sonar) by deliberate operator decision. The operating rule is that the pipeline grades against what a consumer talks to by default, regardless of how a same-day research comparison scored that panel. This is a dated panel change, not a methodology version change. The NO-GO above stands as a finding and is not changed by it.

Appendix: Batch and Cohort IDs

PurposeBatch IDDateRowsLive retrieval
Charleston July baseline (Vol 1 source)mkt_pilot_charleston_2026_07July 7, 2026516Perplexity only
Nashville July baselinecensus_jul2026_nashville_v2July 27-28, 2026984Perplexity only
Charleston Sept pre-shiftcensus_sep2026_charleston_preSept 10, 2026984All four models
Nashville Sept pre-shiftcensus_sep2026_nashville_preSept 10, 2026984All four models
Model test, default tier (ChatGPT, Claude)model_test_free_20260914Sept 14, 2026375 per modelBoth models
Model test, Gemini and Perplexity legsmodel_test_v4_20260914Sept 14, 2026375 per modelBoth models
Model test, flagship (abandoned)model_test_v4_20260914Sept 14, 202684 / 87 (Charleston only)Both models

Retrieval column verified Sept 22, 2026: July worker source (commits 47d45a88, 7a92a389) has no retrieval branch; market_runs.usage_raw is null on all July rows and populated on all September batches. All six batches ran through the real-time market-run route, not the batch API census route first used Sept 18.

Cohorts: charleston_vol1, 687 domains, locked 2026-08-26 19:07:53 UTC. nashville_vol1, 481 domains, locked 2026-09-13 20:54:17 UTC.

Pipeline note for wave 2: MODEL_VERSIONS.perplexity changed from sonar-pro (Chat Completions) to perplexity/sonar (Agent API) on Sept 14, 2026, after every leg in this report had completed. No row in this report ran on the new string.

References

Grittner, T. (2026). The State of AI Recommendations: Charleston 2026. ARO Index. https://doi.org/10.5281/zenodo.21582986. Live at aroindex.com/research/charleston-2026.

Grittner, T. (2026). State of AI Recommendations: Charleston, Vol 2 (July 2026). ARO Index. https://doi.org/10.5281/zenodo.22920524. Live at aroindex.com/research/charleston-2026-vol2.

Cite this report

Grittner, T. (2026). ARO Index Vol 3: The September 2026 Model Shift. ARO Index. https://doi.org/10.5281/zenodo.22921723

Contact

Questions about methodology or data access: reports@tagmakessc.com

ARO(tm), ARO Score(tm), and ARO Index(tm) are trademarks of Therese Grittner. Data current as of Sept 14, 2026. ARO Index publishes live, public research on which local businesses AI models actually select. The audit does not create your score. It reveals it.