Pick a real buyer question from our prompt panel and watch the census work: the same question is asked repeatedly in cold sessions, each answer's recommended brands are extracted, and a league table converges with confidence intervals that narrow as evidence accumulates.
Fraction of sampled answers recommending each brand, with 95% Wilson intervals.
Run the same question twice and the shortlist changes. That variance is why single-run dashboards mislead: a brand can appear in one answer and vanish from the next without anything changing in the market. Sampling turns that noise into an estimate with a stated margin of error, and at benchmark scale (50 prompts, 8 runs, every major engine) the intervals get tight enough to stand behind. See the pilot benchmark for the full-panel view.