Measuring what AI says about you
Visibility measurement Answer-engine governance GEO · AEO
Search used to hand people a set of links and let them choose. Increasingly it hands them a finished answer, assembled from sources they never see. Before anyone optimises for that, someone has to establish what these systems currently say about your organisation — and which sources they draw on to say it. That measurement is the work.
Illustrative, not measured. The point is the shape: the cited sources move between runs of the same prompt, on the same day. One run is an anecdote. A distribution is a finding.
Why this matters now
You no longer choose your sources
A model assembles its answer from whatever it treats as authoritative — an encyclopaedia entry, a news archive, a forum post, a competitor. Your own site is one candidate among many, not the default.
The first impression is second-hand
An investor, a journalist or a prospective hire may form their view of your organisation without ever reaching a page you control. What the model says is the impression.
In regulated sectors this is not a marketing question
An inaccurate machine-generated statement about a product, a study or a financial figure is a compliance matter before it is a traffic matter. It needs an owner, not a dashboard.
How the work runs
Measure first. Decide second. Change third. The order is not a preference — without a baseline there is no way to tell whether a change helped, and no way to defend the decision afterwards. Click any step for the method detail.
Set the scope
One organisation, one audience, one language market.
Method detail
- Mixing audiences produces averages that describe nobody.
- Corporate audiences (investors, media, policy, talent) ask structurally different questions from commercial ones.
- Language market is fixed at the start; results do not transfer between them.
Build the prompt set from real questions
The questions your audiences actually ask — not questions invented to flatter the result.
Method detail
- Drawn from search console data, sales conversations and support enquiries.
- Separated into prompt families: definitional, comparative, procedural, reputational.
- The full question universe is mapped before the tracked subset is chosen, which prevents cherry-picking.
Fix the measurement conditions
Without controls you measure your own browsing history.
Method detail
- Logged out; personalisation and assistant memory disabled.
- Region and interface language defined and held constant.
- Model version and timestamp recorded for every single observation.
Run each prompt repeatedly, across surfaces
Repetition is what turns an observation into a measurement.
Method detail
- Published convergence analysis supports roughly seven runs per prompt per day where brand presence is the question, and around eight where the citation sources matter.
- That is a starting point, not a standard — the right number is the one at which the confidence interval stops narrowing.
- Run across the major assistants and AI answers inside classic search; a non-appearance is itself recorded.
Record what was cited, not just whether you appeared
Presence is the shallow question. Provenance is the useful one.
Method detail
- Every source credited, not only your own; which competitors were named alongside you.
- Whether the statement made about you is factually correct.
- Substitution: does the answer complete the task, so no visit is needed at all?
- For regulated clients, a regulatory-risk flag on any statement that would require review if a human had written it.
Compare against classic search
The single most useful finding this work produces only appears if both are measured.
Method detail
- Ranking well organically while being absent from the generated answer is the gap worth knowing about.
- It usually points at structure and extractability rather than at authority.
- Measured on the same prompt set, same day, same conditions.
Report the distribution, then decide
Findings arrive with their spread and their conditions attached.
Method detail
- No single visibility score. Metrics are reported separately by prompt family.
- Differences smaller than the observed run-to-run variance are reported as noise, not movement.
- Recommendations follow the evidence; where the evidence is thin, the recommendation says so.
Defining the goals
Four questions worth answering before the first measurement. Targets are set with you at the start — a measurement without an agreed objective produces a report nobody can act on.
Presence
Which categories of question should name you at all — and which, honestly, should not?
Not every question is one you need to win.
Source control
What proportion of citations should come from domains you own and can correct?
This is the number governance can actually move.
Accuracy
Which statements about you must be right, and what happens when one is wrong?
Who is notified, who decides, who corrects the source.
Substitution
Where a complete AI answer removes any reason to visit you — is that acceptable for that question?
Sometimes yes. For a directory business, rarely.
What this method cannot tell you
Stated up front, because you will find these out eventually and it is better that you hear them from me.
A finding describes a specific model version on a specific date. It is a snapshot, not a permanent property of your organisation. Re-measurement is part of the work, not an upsell.
Repetition narrows the uncertainty; it never removes it. Small differences between two runs are noise, and I will not present them as movement.
If citations rise after a change, that is correlation. Establishing cause would need controls that live commercial systems almost never permit.
Region, language, account history and personalisation all move the output. Every measurement is conditional on the environment it was taken in, and that environment is documented alongside it.
The work I rely on is recent, largely preprint, and not all of it is disinterested — several of the most-cited papers have authors affiliated with vendors selling measurement tools. I read them, I say which ones, and I do not present their figures as settled.
Nobody can sell you a position inside a generated answer. Anyone offering one is selling something other than what they claim.
Evidence base
Three papers do most of the load-bearing work. Their publication status and their authors' interests are stated, because a method is only as trustworthy as the sources it rests on.
GEO: Generative Engine Optimization
Introduced the term, the GEO-Bench benchmark and the impression metrics still in use across the field.
Status & reading
- Peer-reviewed conference paper; preprint arXiv:2311.09735.
- Its widely quoted visibility gain is a maximum under favourable conditions on a fixed set of candidate sources.
- It is not an average, and it does not establish organic discoverability or durable traffic effects.
Don't Measure Once: Measuring Visibility in AI Search
The basis for treating visibility as a distribution, and for the repetition counts used in step 04.
Status & reading
- arXiv:2604.07585, April 2026, University of St. Gallen. Swiss dataset across four verticals and four engines.
- Preprint — not yet peer reviewed.
- The first author discloses a commercial affiliation. Worth knowing precisely because the paper's conclusion runs against the interest of the measurement-tools industry.
A Critical Survey of Generative Engine Optimization
Reviews the field 2023–2026 and declines to endorse most of what it finds.
Status & reading
- arXiv:2607.14035. Preprint.
- Concludes that no reviewed technique demonstrates a stable, longitudinal, cross-platform causal effect on discoverability or user behaviour.
- This is the source of the position taken on this page: current metrics are tracking indicators, not proven levers.
Figures from this literature are used in client work with their source, date and conditions attached, or not at all. Several widely circulated GEO statistics turn out on checking to have been detached from the paper that produced them.
Wondering what AI currently says about you?
A first engagement is a measurement, not a retainer. A defined prompt set, run under controlled conditions across the major assistants, compared against your classic search performance, and reported with its uncertainty intact. You end up with a baseline you own. What follows from it is your decision, made on evidence rather than on a vendor's promise.
Get in touch