Skip to main content
Enso InsightsEnsoInsights
All posts
September 26, 20267 min read

How AI Visibility Scoring Works: Our Methodology

How Enso Insights scores brand visibility in AI answers: real Brave search snippets in, Gemini primary scoring, GPT-4o cross-check, cited evidence out.

Most "AI visibility" tools ask one model one question and call the result a score. That's a black box: a number with no evidence behind it and no way to verify it.

Our methodology is different, and we publish it because a score you can't audit is a score you shouldn't trust. Here's exactly how an Enso Insights scorecard is produced.

Step 1: Gather real evidence with Brave Search

Before any model scores anything, we collect ground truth. The pipeline calls the Brave Search API (api.search.brave.com/res/v1/llm/context) to pull real, current search snippets about your brand and your competitors — the same kind of retrieval-augmented context that powers AI answers.

Why this matters: we don't ask the models to score from memory alone. The prompt payload they receive includes your brand name, a category hint, your named competitors, and a block of the actual Brave snippets. The models score what the evidence says, and the report shows you that evidence. You're not getting "directional data." You're getting the actual source material next to every score.

Step 2: Gemini scores first (primary scorer)

The full payload goes to the Gemini Developer API (gemini-2.5-flash / gemini-flash-latest) as the primary scorer. Gemini evaluates brand visibility across the dimensions that matter in AI answers: citation presence, competitive position, sentiment of mentions, and the quality of sources being cited.

Gemini is the primary engine because of its strength at structured evaluation over long context blocks — the snippet evidence can be substantial, and the primary scorer needs to hold all of it in working memory while scoring consistently.

Step 3: GPT-4o cross-checks (second scorer)

The same evidence goes to GPT-4o (scoreBrandWithGpt / rankCompetitorsWithGpt) as an independent second scorer. This is the step most tools skip, and it's the whole point of the architecture.

One model scoring alone has a failure mode: its own biases, blind spots, and quirks become your score. A model that has a weak representation of your category, or that happened to overweight one odd snippet, gives you a number that reflects the model as much as your brand. You can't tell the difference — unless a second, independent model scores the same evidence and you compare.

One operational note: the second-engine call is best-effort. If it fails — a missing key, a timeout, or unparseable output — the scorecard renders in clearly-labeled single-engine mode rather than faking a consensus.

Why two models beat one

Independent errors don't correlate. If Gemini misreads a sarcastic review as praise, GPT-4o is unlikely to make the identical mistake on the identical snippet. When both engines converge, you have a genuinely robust signal. When they disagree, you have something even more useful: a flag that says "look here, the evidence is ambiguous."

Consensus, not averaging. We don't just average two scores and call it a day. Averaging buries disagreement — a 90 and a 50 average to 70, which tells you nothing. Our dual-engine consensus preserves the disagreement: you see both engines' read, where they converged, and where the evidence supports two interpretations. Signal survives; noise gets flagged.

Model monoculture is a risk. If your entire visibility strategy is tuned to how one model sees you, you're exposed. Different buyers use different assistants. Scoring with two frontier models from different labs gives you coverage across the two biggest model families your buyers actually use.

Step 4: Cited evidence out

The report shows every score next to the evidence behind it — the actual snippets the models scored. You can read the source, check the model's reading of it, and disagree with us. That's intentional. A methodology that hides its evidence is asking for trust it hasn't earned.

The operational guardrails

A methodology is also what it refuses to do:

  • Fresh reruns are capped at 6 per brand per day on paid plans. AI answers fluctuate run to run; unlimited reruns would let anyone cherry-pick a flattering snapshot. The cap keeps the score honest.
  • Results cache for ~6 hours. Models and search indexes move fast, but not minute to minute. The cache window balances freshness against noise.
  • Cached views are unlimited. Re-reading your report costs nothing and changes nothing. Only new pipeline runs count against the cap.

What the report actually contains

Knowing the pipeline is half the story. Here's what lands in your inbox:

  • Overall visibility score — the dual-engine consensus read on how present and how well-positioned your brand is in AI answers for your category.
  • Competitor comparison — your named competitors scored against the same evidence, so "we're doing fine" becomes "we're cited half as often as X."
  • Sentiment breakdown — not just whether you're mentioned, but what the models say about you when they do: strengths cited, criticisms repeated, gaps noted.
  • Source inventory — the domains and pages the evidence came from, which doubles as your action list. If the models cite a stale third-party page more than your own site, that's the first thing to fix.
  • Engine agreement notes — where Gemini and GPT-4o converged and where they diverged, with the underlying snippets so you can read the disagreement yourself.

Limitations we'll state plainly

No methodology section is honest without these:

  • We measure two model families, not every assistant. Gemini and GPT-4o cover the two biggest labs behind the assistants your buyers use, but Perplexity, Claude, and others have their own retrieval and ranking behavior. Two engines beat one; they don't beat all.
  • Scores are snapshots. A scorecard reflects the evidence and models at run time. Visibility moves as indexes update and models change. That's why the Core plan exists — measurement is a cadence, not an event.
  • Evidence is retrieval, not the model's full memory. The Brave snippet block grounds the scoring in current, verifiable sources. It doesn't capture everything a model "knows" from training. We consider that a feature — training-data impressions can't be audited, but retrieved snippets can.

What we don't claim

We don't claim to know the exact internals of any model's weighting. We don't claim a single score captures everything about your brand's AI presence. And we're careful about data handling: we don't sell your data, and our contracts prohibit training on customer data.

What we do claim is narrow and verifiable: real search evidence in, two independent frontier models scoring it, every score shown next to the evidence that produced it. Read the report, check the snippets, and judge the scoring yourself.

See the methodology on your own brand

The free scorecard runs the full pipeline — the same dual-engine scoring paying customers get. One run, full report, shareable link.

Run your free scorecard →


Written by Enso Insights. Have a question or correction? Email us at support@ensoinsights.us.