Version 0.1 (draft). Published, versioned, and reproducible:
a ruler you can check. Every published figure carries its methodology
version, panel hash, window dates, engine list, and run counts.
1 · What we measure
For a category and a defined buyer-intent prompt panel, per engine and in aggregate:
Share of Recommendation (SoR): probability a brand appears in the recommended set of an answer, estimated by repeated sampling. Primary metric.
Share of Mention: probability of any appearance. Secondary.
Citation Share: probability a page or domain is cited by an answer. The supply-side map.
Retrieval Share: probability a domain is fetched during answer generation, cited or not. Retrieved ⊃ cited; the gap is informative.
2 · Sampling design
Prompt panels of 50+ prompts (200-300 at production), harvested from real buyer language. Panels are versioned and frozen per window; a panel hash is published before results are computed.
8 or more independent runs per prompt per engine per window. Engines are stochastic; single-run measurement is noise presented as signal.
Rolling 21-day windows at production; every observation timestamped.
Cold sessions only: no history, no personalization. We measure the population default and say so.
3 · Extraction
The recommended set is distinguished from incidental mentions. Extraction is
deterministic brand-list matching plus an LLM pass with a published prompt.
Disagreements are logged and human-audited, and a 300-item human-labeled
calibration set gates any extractor change. Citations are harvested from
engines' structured citation output and normalized to domain plus path. All
extractor code and prompts ship with the methodology.
4 · Statistics
Shares are binomial proportions over runs, reported with 95% Wilson score
intervals. No point estimate ships without its interval. Movement claims
require non-overlapping intervals across windows or a two-proportion z-test at
p<0.05, and must survive the market-shift control. Minimum detectable
effects are published so readers know what we cannot see.
5 · Controls
Model updates and index refreshes shift citations for everyone at once. Every
census tracks the full competitive set; per-brand movement is reported relative
to the category baseline, and windows spanning major engine releases are flagged.
6 · Neutrality protocol
No payment contingent on measured results; no placement or optimization services, ever.
Entity lists, panels, and extraction rules pre-registered before computation.
Commercial relationships with measured brands disclosed in the benchmark.
7 · Stated limitations
We measure default, non-personalized sessions; individual answers vary.
Engine APIs may differ from consumer UIs; where we measure API surfaces we say so and validate agreement on a sampled basis.
SoR measures presence in answers, not downstream purchases.
Panels sample buyer language; composition is published so critics can attack it. That is the point.