Methodology

One score hides the truth. We show the runs.

Large language models are non-deterministic: the same prompt, asked twice, returns different answers. Any tool that gives you a single visibility number is averaging away the thing you most need to see. Eidoscore reports ranges, keeps every raw answer, and puts an analyst between the data and your decisions.

Measurement

Repeated runs, honest ranges

Each tracked prompt runs multiple times per engine on a schedule. We report the range and the spread, not a single score, and every data point links to the full stored answer. If the answers disagree, you see that they disagree; a board paper that hides uncertainty does not survive the first person in the room who understands how these models work.

Illustrative example
One prompt, one engine, three runs
Run 1
0
Run 2
0
Run 3
0
Reported range00spread 0
Evidence

What we record for every run

  • Engine and model version
  • Prompt text and prompt version
  • Date and time
  • The full answer
  • Every citation the engine gave
  • Whether the organisation, competitors or the tracked narrative appeared, and how
Sources

Where the answers come from

Every answer we store is labelled: which engine produced it, which model version, which prompt version, the date and time it was captured, and how it was collected. That label travels with the data point, so any figure in your brief can be traced back to the exact answer behind it.

AI answers are not a single fixed thing. They vary between runs, between model versions, between the API and the consumer app, and they can vary by region and language. We treat that as a measurement problem to be stated, not smoothed. A Eidoscore figure describes what the engines returned to our prompts, under recorded conditions, over a recorded period. It is not a claim about what any one person saw on any one day, and we will never present it as one.

Interpretation

How classification works

A language model classifies each item for topic, stance and context: serious, ironic, hostile, supportive, neutral, question, review, misinformation risk, emerging narrative, or lead. Classification quality is measured against a hand-labelled evaluation set, and every insight carries a confidence score and a manual correction field. Misinformation flags are early warnings for a human analyst, never automated accusations, and never claims of coordinated campaigns.

Limits

What we will not do

  • No monitoring of private or closed groups.
  • No profiling of identifiable individuals by political opinion; we monitor topics, narratives and communities.
  • No auto-publishing; nothing leaves the platform without a named human approving it.
  • No claims we cannot evidence; every alert links to its sources, model version and prompt version.

Want to see this on your own narratives?

Request a sample brief