Cited Evidence vs. Black-Box Scores: Why We Show Our Work
Most AI visibility tools hand you a black-box number. We show the real search snippets behind every score — and how our two scoring engines handle disagreement.
Ask most AI visibility tools how they computed your score and you'll get a shrug dressed up as a dashboard: a number, a trend line, maybe a "directional" label. Ask what evidence the number is based on and there's nothing to click. You're expected to trust the box.
We think that's backwards. A score without evidence is an opinion with better marketing. Here's why we built Enso Insights around cited evidence — and what happens when our two scoring engines disagree.
The problem with "directional" scores
"Directional" is doing a lot of work in this industry. It usually means one of three things:
- The tool asked one model one vague question ("How visible is Brand X?") and turned the answer into a number. No retrieved evidence, no documented prompt, no way to reproduce it.
- The score is a proxy, not a measurement. Mentions counted across some corpus, weighted by an undisclosed formula. The number moves and nobody can say why.
- The evidence exists but isn't shown. The tool retrieved something, scored it, and threw away the "something." You get the verdict without the trial.
All three share a failure mode: when the score changes, you can't investigate. Did your visibility actually move, or did the model's mood change? Did a competitor surge, or did the tool tweak its formula? Without the underlying evidence, you're flying blind with extra steps.
What cited evidence means in practice
Every Enso Insights scorecard shows the actual search snippets — pulled live via the Brave Search API — that the scoring engines evaluated. Not summaries. Not "sources include." The snippets themselves, next to the scores they produced.
This changes what you can do with a report:
- Verify the reading. If Gemini scores a snippet as positive and you read it as neutral, you can see exactly where the judgment came from — and push back.
- Find the leverage. The evidence shows which sources the models actually cite when they talk about you: your docs, a review site, a forum thread from 2023. That's your to-do list, ranked by what the machines already trust.
- Track real movement. When your score changes between runs, you can diff the evidence. New press coverage? A competitor's launch? A stale source dropping out? The cause is visible instead of mysterious.
Evidence doesn't just make the score trustworthy. It makes it actionable. A number tells you where you stand. The snippets tell you what to do next.
The disagreement problem (and why it's a feature)
Run two independent frontier models over the same evidence and they will sometimes disagree. Most dual-model systems handle this by averaging — which is the worst option. A 90 and a 50 don't mean 70. They mean the evidence supports two genuinely different readings, and averaging erases the most interesting fact in the report.
We handle disagreement with consensus, not averaging:
- Convergence is the strong signal. When Gemini (primary scorer) and GPT-4o (second scorer) independently land in the same place, that score has survived two different architectures, two different training corpora, two different sets of biases. That's as close to ground truth as this field gets.
- Divergence is flagged, not hidden. When the engines disagree, the report says so — and shows you the evidence both of them read. Often the disagreement itself is the insight: a sarcastic review, a brand name shared with an unrelated company, a category where the models genuinely have different training exposure. That's not noise to smooth over. That's the thing you most need to know, because your buyers' assistants will disagree with each other too.
- No engine gets a veto. The primary scorer doesn't overrule the second. The report preserves both reads so you can see the shape of the uncertainty instead of a false point estimate.
This is the honest way to use two models. Averaging pretends uncertainty doesn't exist. Consensus reporting shows you exactly where it lives.
Why this matters more than it sounds
Your buyers don't all use the same assistant. Some ask ChatGPT, some ask Gemini, some ask Perplexity. If the models disagree about your brand — and they will, on the margins — a single-model score is a coin flip dressed as precision.
A methodology that surfaces disagreement is the only one that reflects the world your buyers actually inhabit: a world where the answer depends on which machine they asked. We'd rather show you the messy truth than a clean fiction.
The bar we'd hold competitors to
We don't name competitors here — the space is young and most tools are iterating fast. But we'd suggest holding any AI visibility tool to three questions before you pay for it:
- Show me the evidence. Can I see the actual sources behind my score, or just the score?
- How many models? One model's read is one model's biases. What happens when a second model disagrees?
- Can I reproduce it? If I run the same audit tomorrow, will I get the same evidence trail — or a different black box?
If a tool can't answer all three, you're buying a number, not a measurement.
Get a scorecard you can actually audit
The free scorecard runs the full pipeline: real Brave search evidence, Gemini primary scoring, GPT-4o cross-check, every score next to its cited evidence. Judge our methodology against the evidence yourself.
Written by Enso Insights. Have a question or correction? Email us at support@ensoinsights.us.