How to Measure AI Visibility: Metrics, Methods and Benchmarks
To measure AI visibility, ask each AI engine the questions your buyers ask, on a schedule, in your market, and count how often the answers recommend you. Track visibility (the share of answers that recommend you), share of voice, per-question wins over 28 days, citations of your own domain and visits from AI. Read the numbers as a trend over weeks, never as a single screenshot.
Key takeaways
- One answer proves nothing: the same question can name different brands on the next run, so measure shares across many answers.
- Lead with visibility and share of voice. Average rank mostly measures how often engines reshuffle their lists.
- A fixed set of buyer questions, asked per engine on a fixed cadence, is what makes week-over-week changes meaningful.
- In Citepoint's October 2026 index, each category's leading brand was named in a median of 19 of 20 buyer questions, and the fifth brand in 10.
What is AI visibility?
AI visibility is how often AI assistants recommend your brand when people ask them the questions that lead to a purchase. It's measured per engine (ChatGPT, Gemini, Claude, Perplexity, Google AI Overviews and AI Mode), per question and over time, because each engine searches the web differently and each answer can change from one run to the next.
Why isn't one screenshot a measurement?
Ask the same question twice and you may get two different shortlists. Engines re-run their web searches, read different pages and reword their answers, and logged-in users carry memory and context the engine weighs too. A screenshot shows one sample. A measurement counts many samples, from a consistent setup, and compares periods.
Which metrics should you track?
| Metric | What it counts | What it tells you |
|---|---|---|
| Visibility | The share of answers in a period that recommend you | How often buyers hear your name |
| Share of voice | Your mentions as a share of all brand mentions in those answers | How you compare with competitors |
| Questions won | For each question, how many of its last 28 days of answers recommended you | Which questions you own and which you're losing |
| Cited as a source | How often engines link to your own domain | Whether your pages are being read and trusted |
| AI referrals | Visits from chatgpt.com, perplexity.ai, gemini.google.com and other assistants | Whether visibility turns into traffic |
Compare each metric with the period before it, for example the last 7 days against the 7 days before, so a single noisy day doesn't look like a trend.
Why not lead with average rank?
Because AI answers rarely list products in the same order twice. When a brand is named, it might be second today and fourth tomorrow for no reason that you can act on. Average position mostly measures that shuffle. Whether you're on the list at all is the signal that moves with your work. That's why Citepoint's methodology leads with how often you're recommended, not where.
How do you build the question set?
Start with the questions that lead to a purchase, not the ones that lead to a definition:
- Best for: "best help desk for a five-person support team"
- Alternatives: "alternatives to Zendesk for small teams"
- Comparisons: "Freshdesk vs Help Scout"
- Problems: "how do I stop tickets getting lost between inboxes"
Write 20 to 50 of them, in the words your buyers use, and keep the set fixed. Adding and removing questions every week makes visibility rise and fall for reasons that have nothing to do with your brand. The free Buyer Prompt Generator drafts a starting set.
How do you collect the answers?
Four choices decide whether your numbers mean anything:
- Where the answer comes from. Buyers use the consumer apps, which search the web and show sources. An engine's API without web search answers from training data alone and will lag behind. Ask the apps where you can, or the API with live web search on.
- Market. Answers differ by country and language. Ask in the market you sell to.
- Cadence. Daily for engines your buyers use most, weekly for the rest. Keep it fixed.
- A fixed model, where you choose one. If the model changes mid-month, your trend line changes with it.
Then read each answer the same way every time: list every product it recommends or presents as an option, in order, with its domain, and match the list against your brand, its aliases and your competitors. If the engine gave no answer, for example when Google showed no AI Overview, count that answer neither for nor against you.
What does good look like?
It depends on your category and your stage, but there are useful reference points. In Citepoint's AI Recommendation Index for October 2026, we asked 20 buyer questions in each of 16 software categories on ChatGPT, Gemini and Google AI Overviews (US, English). Counting a brand when at least one of the three engines named it:
| Position in the category | Questions named in, median across 16 categories | Range |
|---|---|---|
| Most-named brand | 19 of 20 | 16 to 20 |
| Fifth most-named brand | 10 of 20 | 5 to 13 |
So the brands on a category's shortlist show up in at least half its buyer questions, and the leader in nearly all of them. A newer brand usually starts near zero. The goal is movement: more questions won this month than last, on more engines.
How do you measure the traffic it brings?
AI assistants send visitors with recognizable referrers, such as chatgpt.com, perplexity.ai and gemini.google.com. Group them into one channel in your analytics so you can see visits, sign-ups and revenue from AI separately. The free AI traffic setup for GA4 builds the channel group for you. Expect traffic to lag behind visibility: buyers often read the answer, then search your name later.
What are the common mistakes?
- Changing the question set every week, which breaks the trend.
- Mixing engines into one number too early. A brand can lead on Gemini and be missing on ChatGPT.
- Testing in a logged-in account full of your own history.
- Reacting to one day. Read the numbers as a direction over weeks.
- Measuring without acting. The point of tracking is to find the questions you're losing, publish the page that answers them and earn mentions on the sources engines cite, then see what moved.
You can run all of this in a spreadsheet: a fixed list of questions, one row per engine per day, and a column for each brand named. Citepoint does it across six engines on a schedule and connects the losses to the pages that win them back. The free scan shows where you stand on three of your buyer questions in a minute.
Frequently asked questions
How many questions do I need to measure AI visibility?
Twenty is enough to see a trend in one category, and 50 covers most companies' main use cases. What matters more than the number is keeping the set fixed so periods can be compared.
How often should I check?
Daily for the engines your buyers use most and weekly for the rest. Compare rolling periods, such as the last 7 days against the 7 before, rather than single days.
Should I use an engine's API or its app?
The app is what buyers see. If you use an API, turn web search on, because an API answering from training data alone won't reflect recent pages or mentions.
Is a citation the same as a recommendation?
No. An engine can cite your page as a source while recommending a competitor, or recommend you without linking to your site. Track both, separately.