One answer is noise.
Eight answers are a measurement.

Pick a real buyer question from our prompt panel and watch the census work: the same question is asked repeatedly in cold sessions, each answer's recommended brands are extracted, and a league table converges with confidence intervals that narrow as evidence accumulates.

Pilot data. This demo replays simulated engine answers so the mechanism is visible without live inference. Production benchmarks sample real AI engines the same way.

Sampled answer

run 0 of 8
Press "Run the census" to start sampling this question.

Share of Recommendation

Fraction of sampled answers recommending each brand, with 95% Wilson intervals.

What this shows

Run the same question twice and the shortlist changes. That variance is why single-run dashboards mislead: a brand can appear in one answer and vanish from the next without anything changing in the market. Sampling turns that noise into an estimate with a stated margin of error, and at benchmark scale (50 prompts, 8 runs, every major engine) the intervals get tight enough to stand behind. See the pilot benchmark for the full-panel view.