Research
How to Measure AI Visibility
What can be measured today, what cannot, and how to build a defensible baseline.
- Author
- By Anne Mason
- Published
- Published September 6, 2026
- Updated
- Updated September 6, 2026
- Reading time
- 7 min read
Anyone selling AI visibility measurement should begin with an admission: nobody outside the model providers can see the full picture. There is no console that reports how often an assistant mentioned your company last month, to whom, or what it said. Anyone implying otherwise is selling confidence they do not have.
What follows is what can honestly be measured today, what cannot, and how to build a baseline you would be comfortable defending to a sceptical board.
What can be measured
Your own machine-readability. Whether pages fetch cleanly and quickly, whether robots directives block anything important, whether canonical tags and sitemaps are sane, whether structured data exists and validates, whether an llms.txt file is present. These are binary or near-binary checks and they are entirely within your control.
Your entity clarity. Whether the site states plainly what the organisation is, does and serves; whether Organization schema exists with consistent identifiers; whether your name and category are used consistently across your profiles. Partly automatable, partly a judgement call, but reproducible.
Model recognition. Ask a set of models, in a controlled way, whether they know your company and how they describe it. Record whether the description is accurate, which sources they appear to draw on, and how much the models agree with each other. Agreement is a useful proxy for how settled your entity is.
Buyer-question coverage. Define a fixed set of questions a real buyer would ask before knowing any vendor names. Ask them across multiple models on a schedule. Record whether you are named, in what position, alongside whom, and with what characterisation. Track the change over time.
Competitor presence. The same exercise, recording who does appear. This is often the most sobering output, because it shows the consideration set you are competing against and how it is described.
What cannot be measured
Actual query volume. You do not know how many buyers asked. You only know what happens when you ask.
Actual influence. You cannot see a buyer read an answer, absorb a name and search for it a week later. The causal chain from AI mention to pipeline is real but mostly invisible in analytics.
Training data composition. You cannot inspect what a model learned about you or when. You can only infer it from behaviour.
Stability. Model outputs vary between runs, versions and phrasing. A single query result is an anecdote. Only repeated, structured sampling produces something like a signal.
How to build a defensible baseline
Start with a fixed question set. Fifteen to thirty buyer questions, written in natural language, covering the ways your market actually phrases its problems. Do not include your own brand name; the point is to test unprompted recall.
Fix the models. Choose a small panel of systems and keep it stable so change over time means something.
Sample repeatedly. Ask each question more than once and treat the aggregate as the result, not any single answer.
Record structure, not impressions. For each answer: named or not, position, competitors named, accuracy of description, apparent sources. Store it so it can be compared.
Separate measured from estimated. Some things you will know. Some you will infer. Label them differently and never let an estimate borrow the credibility of a measurement.
Repeat on a schedule. Monthly is a reasonable cadence. The value is in the trend, not the snapshot.
A number you cannot reproduce is not a measurement. It is a mood.
Why we built it this way
Beacon is Market Arc's implementation of exactly this discipline. It reads a site the way a machine does, checks the foundations and the entity signals, asks buyer questions across a panel of models, records recognition and agreement, and scores the result across eight weighted pillars. Where a pillar depends on data we cannot yet access, the score is labelled as an estimate and the limitation is shown rather than hidden.
The result is not a perfect view of AI visibility. Nobody has one. It is a repeatable, honest baseline that lets a company see where it stands, what is actually broken, and whether the work it does next moves the numbers. That is the standard we would want applied to us, and it is the standard we hold Beacon to.