Searchestrablog

How Searchestra measures AI visibility

AI visibility is only as trustworthy as the method behind it. This page explains how we measure, in plain terms, so you can judge our numbers rather than take them on faith. It is deliberately specific about what we do, and honest about what we do not claim.

Why we publish our method

Two AI visibility tools can measure the same brand and report different numbers, because they use different questions, engine coverage and definitions. That is exactly why the method matters more than the headline score. We would rather you trust a number you can interrogate than a confident figure you have to take on faith. This mirrors the transparency the IAB's Measuring Visibility in the AI Era framework asks measurement providers to compete on. See also directional vs decision-grade and what to ask a provider.

The AI answer engines we track

We maintain a catalog of AI answer engines spanning assistants and AI-search surfaces. Each is measured on its own, never blended into a single vanity score:

EngineProviderType
ChatGPTOpenAIAssistant
PerplexityPerplexityAI search
Google AI OverviewsGoogleAI search
Google AI ModeGoogleAI search
GeminiGoogleAssistant
ClaudeAnthropicAssistant
Microsoft CopilotMicrosoftAssistant
GrokxAIAssistant
DeepSeekDeepSeekAssistant

Honest caveat: an engine is only queried when access to it is configured. Coverage for a given account depends on which engines are enabled, so we report per-engine rather than implying every surface is always live. See multi-platform aggregation.

The unit of measurement

Our unit is not a single API call but a question paired with an engine, run as a real request, not a simulation. Results are reported per engine so a strong showing on one surface never hides invisibility on another. This is the difference between share of model and share of voice: we keep the layers separate rather than averaging them into one blurry figure.

Where the questions come from

The question set for a category is generated from a taxonomy of topics and a graph of realistic buyer scenarios, rather than hand-picked, so coverage is systematic instead of cherry-picked toward questions a brand happens to win. Each question carries its own metadata (intent, scenario, source type, version).

On sourcing, which the IAB framework asks providers to be explicit about: the universe combines the taxonomy and scenarios with evidence of real search demand, and uses synthetic expansion only to fill coverage gaps, never as the whole set. Search-demand signals inform the universe but do not solely define it, and every question records which source type it came from.

Questions are classified by intent: informational, commercial (including recommendation-style "best X" queries), transactional, navigational and comparison. This spans the intent types the IAB decision-grade approach asks to be covered and segmented, and the classifier runs in both English and Turkish, so a clearly commercial Turkish question is not silently treated as a general one.

The question set is brand-neutral by construction: the brand being measured does not influence which questions are generated. Brand identity is used only as a downstream accept-or-reject check on already-generated candidates, never to shape what is asked. That separation is what keeps the set from being quietly tilted toward a flattering result, the guardrail behind prompt and platform bias.

The set is versioned and immutable: a measurement run records exactly which version of the question set it used, so a change in results can be told apart from a change in questions. See measurement stability and reproducibility.

For each project we can disclose the transparency profile of the exact question-set version in use: how many questions it holds, how they break down by intent and by source type, the weighting basis, and the baseline and refresh policy behind it. These cover several of the query-set transparency dimensions the IAB framework identifies, so the question set behind a score is not a black box; we are still expanding what we expose here.

The four layers we report

We report along the four layers of the IAB framework rather than one score, so you can see where representation breaks down. The layer names are the IAB framework's; the example metrics in the third column are how Searchestra operationalizes each layer:

IAB layerThe question it answersHow Searchestra reports it (examples)
PresenceAre you in the answer at all?mention rate, share of voice
ProminenceHow central are you when present?average rank, citation share
PortrayalHow are you described?sentiment
PersuasionAre you actually recommended?recommendation share

These are transparent, separately reported measures, not a black-box composite (some are ratios, some like average rank and sentiment are not). A note on naming: only the three formulas below carry the IAB framework's exact definition. Our operational metrics are not always the IAB metric of a similar name, for example our citation share (a brand's share of all citations) is not the IAB Citation Rate (cited responses over total responses), and our recommendation share is not the IAB Recommendation Strength scale. Recommendation share counts only answers where you are genuinely recommended as a first choice, kept separate from being merely mentioned. See the 4 P's of AI visibility, mention rate and citation rate.

The formulas, in the open

Where a metric matches the definition in the IAB framework, we compute it exactly that way, no proprietary reweighting. These are the ones that carry a shared, standardized formula:

MetricHow it is calculated
Mention RateResponses that mention your brand, divided by total responses in the query set.
Share of VoiceYour brand mentions as a proportion of all tracked brand mentions in the competitive set.
Visibility MomentumThe percentage change in Mention Rate or Share of Voice between two measurement periods.

Share of Voice depends on two choices the IAB framework requires a provider to make explicit, because they change the number. Here is how we make them: the competitive category is the set of brands tracked for your project (defined at setup and visible to you), and the total mention universe is sized as the sum of mentions across exactly that tracked set, so your share is your mentions over that total, nothing hidden in the denominator. Two tools can report different Share of Voice for the same brand purely from these two choices, which is why we surface ours. Momentum is read carefully too: when a platform change shifts the baseline, we re-baseline and say so, rather than book it as a brand win.

Reproducibility and change

Because AI answers vary between runs, a single answer proves nothing. We measure on the versioned question set and aggregate across runs so one volatile response does not masquerade as a trend. When a platform changes, we treat it as a new baseline event, not a marketing win. See visibility momentum and understanding AI non-determinism.

The honesty principle

We report what we collect and flag what we cannot. When a signal is not measured, for example if recommendation judging is not enabled for a run, the field is left empty rather than filled with a guess. We do not present a confident single number where the data only supports a directional read. A documented limitation is part of the rigor, not a weakness to hide.

What we do not claim

An API result is not identical to what a signed-in person sees in a product's interface, and we are explicit about which layer a given number reflects. Some signals, such as sentiment, combine a deterministic lexicon with a model-based read and are treated as signals rather than verdicts. Where a number is directional, we say so plainly. See directional vs decision-grade measurement.

Key takeaway.

We measure AI visibility across a catalog of answer engines using a brand-neutral, versioned question set, report the four IAB layers as separate, transparent measures rather than one blended score, and flag what we cannot measure rather than guess.

Frequently asked questions

Do you publish your exact prompts?

No. The question set is versioned and brand-neutral, and we disclose the method and its metadata rather than the proprietary set. Publishing the exact prompts would let them be gamed, which would corrupt the measurement for everyone.

Why do two AI visibility tools disagree?

Different question sets, engine coverage, competitive sets and definitions. Two tools can measure the same brand and legitimately differ, which is exactly why a disclosed method matters more than the headline number.

Is this decision-grade measurement?

We build toward the IAB decision-grade criteria: versioned questions, aggregation across runs, per-engine reporting and brand-neutral construction. Where a metric can only support a directional read, we label it that way rather than dress it up.