Measuring what AI says about you

Visibility measurement Answer-engine governance GEO · AEO

Search used to hand people a set of links and let them choose. Increasingly it hands them a finished answer, assembled from sources they never see. Before anyone optimises for that, someone has to establish what these systems currently say about your organisation — and which sources they draw on to say it. That measurement is the work.

One prompt, seven runs, one day Sources cited per run
Run 01WikipediaNews archiveyour-domain.com
Run 02WikipediaAnalyst blogCompany register
Run 03News archiveTrade pressWikipedia
Run 04your-domain.comWikipediaCompetitor site
Run 05WikipediaForum threadTrade press
Run 06Wikipediayour-domain.comNews archive
Run 07Competitor siteTrade pressCompany register

Illustrative, not measured. The point is the shape: the cited sources move between runs of the same prompt, on the same day. One run is an anecdote. A distribution is a finding.

Why this matters now

The mechanism 01

You no longer choose your sources

A model assembles its answer from whatever it treats as authoritative — an encyclopaedia entry, a news archive, a forum post, a competitor. Your own site is one candidate among many, not the default.

The exposure 02

The first impression is second-hand

An investor, a journalist or a prospective hire may form their view of your organisation without ever reaching a page you control. What the model says is the impression.

The stakes 03

In regulated sectors this is not a marketing question

An inaccurate machine-generated statement about a product, a study or a financial figure is a compliance matter before it is a traffic matter. It needs an owner, not a dashboard.

How the work runs

Measure first. Decide second. Change third. The order is not a preference — without a baseline there is no way to tell whether a change helped, and no way to defend the decision afterwards. Click any step for the method detail.

Step 01 · Framing Scope

Set the scope

One organisation, one audience, one language market.

Method detail

  • Mixing audiences produces averages that describe nobody.
  • Corporate audiences (investors, media, policy, talent) ask structurally different questions from commercial ones.
  • Language market is fixed at the start; results do not transfer between them.
Audience definition
Step 02 · Inputs Prompts

Build the prompt set from real questions

The questions your audiences actually ask — not questions invented to flatter the result.

Method detail

  • Drawn from search console data, sales conversations and support enquiries.
  • Separated into prompt families: definitional, comparative, procedural, reputational.
  • The full question universe is mapped before the tracked subset is chosen, which prevents cherry-picking.
Prompt familiesSearch console
Step 03 · Controls Conditions

Fix the measurement conditions

Without controls you measure your own browsing history.

Method detail

  • Logged out; personalisation and assistant memory disabled.
  • Region and interface language defined and held constant.
  • Model version and timestamp recorded for every single observation.
Reproducibility
Step 04 · Sampling Repetition

Run each prompt repeatedly, across surfaces

Repetition is what turns an observation into a measurement.

Method detail

  • Published convergence analysis supports roughly seven runs per prompt per day where brand presence is the question, and around eight where the citation sources matter.
  • That is a starting point, not a standard — the right number is the one at which the confidence interval stops narrowing.
  • Run across the major assistants and AI answers inside classic search; a non-appearance is itself recorded.
~7–8 runsMulti-engine
Step 05 · Capture Citations

Record what was cited, not just whether you appeared

Presence is the shallow question. Provenance is the useful one.

Method detail

  • Every source credited, not only your own; which competitors were named alongside you.
  • Whether the statement made about you is factually correct.
  • Substitution: does the answer complete the task, so no visit is needed at all?
  • For regulated clients, a regulatory-risk flag on any statement that would require review if a human had written it.
Source provenanceRisk flag
Step 06 · Baseline Comparison

Compare against classic search

The single most useful finding this work produces only appears if both are measured.

Method detail

  • Ranking well organically while being absent from the generated answer is the gap worth knowing about.
  • It usually points at structure and extractability rather than at authority.
  • Measured on the same prompt set, same day, same conditions.
Organic vs generated
Step 07 · Output Reporting

Report the distribution, then decide

Findings arrive with their spread and their conditions attached.

Method detail

  • No single visibility score. Metrics are reported separately by prompt family.
  • Differences smaller than the observed run-to-run variance are reported as noise, not movement.
  • Recommendations follow the evidence; where the evidence is thin, the recommendation says so.
Distribution, not score

Defining the goals

Four questions worth answering before the first measurement. Targets are set with you at the start — a measurement without an agreed objective produces a report nobody can act on.

Goal

Presence

Which categories of question should name you at all — and which, honestly, should not?

Not every question is one you need to win.

Goal

Source control

What proportion of citations should come from domains you own and can correct?

This is the number governance can actually move.

Goal

Accuracy

Which statements about you must be right, and what happens when one is wrong?

Who is notified, who decides, who corrects the source.

Goal

Substitution

Where a complete AI answer removes any reason to visit you — is that acceptable for that question?

Sometimes yes. For a directory business, rarely.

What this method cannot tell you

Stated up front, because you will find these out eventually and it is better that you hear them from me.

Model drift

A finding describes a specific model version on a specific date. It is a snapshot, not a permanent property of your organisation. Re-measurement is part of the work, not an upsell.

Irreducible variance

Repetition narrows the uncertainty; it never removes it. Small differences between two runs are noise, and I will not present them as movement.

No causal proof

If citations rise after a change, that is correlation. Establishing cause would need controls that live commercial systems almost never permit.

Conditional results

Region, language, account history and personalisation all move the output. Every measurement is conditional on the environment it was taken in, and that environment is documented alongside it.

Young research base

The work I rely on is recent, largely preprint, and not all of it is disinterested — several of the most-cited papers have authors affiliated with vendors selling measurement tools. I read them, I say which ones, and I do not present their figures as settled.

No guaranteed placement

Nobody can sell you a position inside a generated answer. Anyone offering one is selling something other than what they claim.

Evidence base

Three papers do most of the load-bearing work. Their publication status and their authors' interests are stated, because a method is only as trustworthy as the sources it rests on.

Aggarwal et al. · ACM SIGKDD 2024 Foundation

GEO: Generative Engine Optimization

Introduced the term, the GEO-Bench benchmark and the impression metrics still in use across the field.

Status & reading

  • Peer-reviewed conference paper; preprint arXiv:2311.09735.
  • Its widely quoted visibility gain is a maximum under favourable conditions on a fixed set of candidate sources.
  • It is not an average, and it does not establish organic discoverability or durable traffic effects.
Peer-reviewedGEO-Bench
Schulte, Bleeker & Kaufmann · 2026 Sampling

Don't Measure Once: Measuring Visibility in AI Search

The basis for treating visibility as a distribution, and for the repetition counts used in step 04.

Status & reading

  • arXiv:2604.07585, April 2026, University of St. Gallen. Swiss dataset across four verticals and four engines.
  • Preprint — not yet peer reviewed.
  • The first author discloses a commercial affiliation. Worth knowing precisely because the paper's conclusion runs against the interest of the measurement-tools industry.
PreprintInterest disclosed
Martinez · 2026 Survey

A Critical Survey of Generative Engine Optimization

Reviews the field 2023–2026 and declines to endorse most of what it finds.

Status & reading

  • arXiv:2607.14035. Preprint.
  • Concludes that no reviewed technique demonstrates a stable, longitudinal, cross-platform causal effect on discoverability or user behaviour.
  • This is the source of the position taken on this page: current metrics are tracking indicators, not proven levers.
No proven causality

Figures from this literature are used in client work with their source, date and conditions attached, or not at all. Several widely circulated GEO statistics turn out on checking to have been detached from the paper that produced them.

Wondering what AI currently says about you?

A first engagement is a measurement, not a retainer. A defined prompt set, run under controlled conditions across the major assistants, compared against your classic search performance, and reported with its uncertainty intact. You end up with a baseline you own. What follows from it is your decision, made on evidence rather than on a vendor's promise.

Get in touch