CLIRO DATASET

What AI engines actually cite,measured every day

Cliro runs prompts against the major AI engines every day and records which brands get mentioned and which sources get cited. This page explains what that dataset measures, how it is collected and what it does not prove.

  • Collected from real answers, not from model self-reports
  • Same collection method across every engine
  • Method and limitations published alongside the numbers

At a glance

AI engines covered
7
Collection cadence
Daily

Only figures we have actually measured are listed here. Anything not yet published is left out rather than estimated.

What is in the dataset

Every record ties one answer from one engine to the brands mentioned and the sources cited in it. Coverage spans ChatGPT, Gemini, Claude, Perplexity, Grok, Google AI Mode and Microsoft Copilot.

By engine

Each answer is attributed to the engine that produced it, so coverage can be compared engine by engine instead of averaged into a single number.

By category

Prompts are grouped by the industry and buying intent they belong to, which is what makes visibility comparable between brands that compete for the same answers.

By market

Answers are collected per market, so a brand that is visible in one country and invisible in another does not average out into a misleading middle.

By cited source

Every citation keeps the domain it came from, which is what turns visibility into something actionable: you can see which sites the models actually read.

Over time

Collection is continuous, so changes can be read as a trend rather than as a single snapshot that may or may not be representative.

How it is collected

The whole method in four steps. If a number on this page cannot be traced back through them, it does not belong here.

  1. 1

    Prompts are executed, not scraped

    Each prompt is sent to each engine and the answer is stored as served. Nothing is inferred from what a model says about itself.

  2. 2

    Mentions and citations are separated

    A brand named in the text and a source linked as a citation are recorded as different things, because they are: being mentioned and being the source are not the same kind of visibility.

  3. 3

    Aggregation happens after attribution

    Rates are always computed over total answers, never over mentions, so a brand cannot look better simply by appearing in fewer answers.

  4. 4

    Collection is continuous

    Prompts re-run on a fixed cadence, which is what makes the series comparable across time instead of a set of unrelated snapshots.

What this dataset does not prove

Stated up front, because the limits are what make the rest worth trusting.

  • It is not a census of AI answers

    It covers the prompts that are tracked, not everything anyone asks. It describes a measured sample, and should be read as one.

  • Results depend on the prompts

    Change the wording and the answer changes. Comparisons only hold between brands measured on the same prompt set.

  • Models change underneath the data

    Engines update their models without notice, so a shift in visibility can come from the model rather than from anything the brand did.

Citing this dataset

Use of the figures is free, including commercially. Attribution with a link is all we ask, and it is what lets a reader verify the number.

Suggested citation

Cliro Dataset, Cliro — https://www.usecliro.com/cliro-dataset