AI Answer Measurement Standard

Version 0.1 (draft). Published, versioned, and reproducible: a ruler you can check. Every published figure carries its methodology version, panel hash, window dates, engine list, and run counts.

1 · What we measure

For a category and a defined buyer-intent prompt panel, per engine and in aggregate:

2 · Sampling design

3 · Extraction

The recommended set is distinguished from incidental mentions. Extraction is deterministic brand-list matching plus an LLM pass with a published prompt. Disagreements are logged and human-audited, and a 300-item human-labeled calibration set gates any extractor change. Citations are harvested from engines' structured citation output and normalized to domain plus path. All extractor code and prompts ship with the methodology.

4 · Statistics

Shares are binomial proportions over runs, reported with 95% Wilson score intervals. No point estimate ships without its interval. Movement claims require non-overlapping intervals across windows or a two-proportion z-test at p<0.05, and must survive the market-shift control. Minimum detectable effects are published so readers know what we cannot see.

5 · Controls

Model updates and index refreshes shift citations for everyone at once. Every census tracks the full competitive set; per-brand movement is reported relative to the category baseline, and windows spanning major engine releases are flagged.

6 · Neutrality protocol

7 · Stated limitations