How to measure your brand's visibility in ChatGPT, Perplexity and Google AI Overviews
Your usual analytics are blind to AI answers. Here's how to measure your brand's visibility across ChatGPT, Perplexity and Google AI Overviews - the metrics that matter, a copy-paste prompt matrix, and the common mistakes to avoid.

How to measure your brand's visibility in ChatGPT, Perplexity and Google AI Overviews
A growing share of your customers now form their first impression of your brand inside an AI answer, one they get from ChatGPT, Perplexity or Google's AI Overviews. The problem is that your usual tools are blind to it. Google Analytics doesn't track it. Your SEO platform doesn't track it. Social listening definitely doesn't. And the moment that matters most is silent: if your brand isn't mentioned, nothing gets logged. No click, no impression, no bounce rate. The opportunity just evaporates.
The good news is that AI visibility is measurable. This guide shows you exactly how, including a copy-paste prompt matrix you can run this afternoon. (If you want the bigger picture first, start with our complete guide to Generative Engine Optimization.)
What "visibility" actually means in an AI answer
Before you measure anything, agree on what you're measuring. Five metrics carry the weight:
Mention rate: the percentage of tested prompts where your brand appears at all. This is the headline number.
Share of voice (or Share of Model): how often you appear versus named competitors on the same prompts. This is your competitive benchmark.
Citation rate: whether the answer actually links to you, not just names you. Citations drive trackable traffic; mentions build association.
Position and prominence: where you land in the answer. AI engines tend to treat the first-named brand as the default recommendation, so being buried in the last paragraph is not the same as leading the list.
Sentiment and accuracy: not just whether you appear, but how you're described. A confident but outdated mention (wrong pricing, wrong positioning) can cost you the deal.
Why you have to measure each engine separately
Here's the mistake almost everyone makes: treating "AI visibility" as one number. It isn't. ChatGPT, Perplexity and Google AI Overviews pull from different sources and cite in completely different ways, so a strong score in one tells you almost nothing about the others. You have to read each on its own terms.
ChatGPT mentions brands generously but links to them rarely. By some estimates only around 20% of ChatGPT mentions include a clickable citation, meaning roughly 80% of the recommendations and comparisons shaping opinions are invisible to traditional analytics. It uses retrieval-augmented generation with "fan-out" sub-queries, and leans heavily on encyclopedic and community sources like Wikipedia, Reddit, Amazon and Forbes.
Perplexity is the opposite. It crawls the web in real time and almost always includes clickable inline citations (typically four to eight sources per answer), which makes it the one major engine where AI visibility translates directly into trackable referral traffic in GA4. It's also Reddit-heavy: if you're not present in the relevant community threads, you're often not present in Perplexity.
Google AI Overviews now appear on roughly 48% of all searches, up from about 34.5% at the end of 2025. They show a strong brand preference (close to 60% of citations go to brand domains, per Otterly's 2026 data), but they only surface for a share of queries, and most AI search sessions end without a click. That means the mention itself often is the entire impression.
Engine | Citation behavior | Trackable traffic? | Leans on |
|---|---|---|---|
ChatGPT | Names brands, links ~20% of the time | Mostly no | Wikipedia, Reddit, Amazon, Forbes |
Perplexity | Always cites, 4-8 linked sources | Yes (GA4 referrals) | Reddit, community content, live web |
Google AI Overviews | Brand-leaning citations, ~60% to brand domains | Rarely (mostly no-click) | YouTube, Wikipedia, Forbes, Quora |
The manual method: a spot check you can run this afternoon
You don't need a platform to get a directional baseline. You need the right prompts. Rather than guess which questions to test, copy the matrix below, swap the bracketed placeholders for your own details, and run each prompt in all three engines.
Placeholders: [category] = your product/service category · [brand] = your brand · [competitor] = a key rival · [role] = your buyer's job title · [problem] = the problem you solve · [region] = your market
Prompt type | What it reveals | Copy-paste template |
|---|---|---|
Category / discovery | Whether you're in the consideration set at all | "What are the best [category] for [use case]?" · "Which [category] providers should I consider in [region]?" |
Comparison | Your positioning against a named rival | "How does [brand] compare to [competitor]?" · "[brand] vs [competitor]: which is better for [use case]?" |
Persona / use-case | Whether you're recommended in your buyer's context | "What should a [role] look for when choosing a [category]?" · "Best [category] for a [role] at a mid-sized company" |
Problem / jobs-to-be-done | Top-of-funnel presence before your category is named | "How do I [problem]?" · "What's the best way to solve [problem]?" |
Brand-direct / reputation | Accuracy and sentiment: how you're described | "What is [brand]?" · "Is [brand] any good?" · "What are the pros and cons of [brand]?" |
A worked example:
Category: "What are the best AI visibility platforms for European brands?" Comparison: "How does 3RD compare to Profound?" Persona: "What should a CMO look for when choosing an AI visibility tool?" Problem: "How do I find out if ChatGPT recommends my brand?" Brand-direct: "What is 3RD and who is it for?"
How to run it: Test 10–15 prompts across the five types. Run each one three to five times per engine, because answers vary between runs. For every response, log five things: whether your brand appears, its position, which competitors show up, whether you're linked or just named, and whether the description is accurate. In an afternoon you'll have a real baseline, and usually a few uncomfortable surprises.
The limit of doing this by hand is volume. Twenty-five prompts run five times each is 125 data points: enough to see obvious patterns, not enough for statistically stable visibility rates. That's the point where a manual spot check turns into continuous, automated tracking.
Common mistakes when measuring AI visibility
Measuring AI visibility is easy to get wrong in ways that make the numbers look precise but mean little. Watch for these:
Chasing rank position instead of visibility rate. In an AI answer, position swings wildly from run to run. What's stable, and what matters, is how often you appear across many runs. As SparkToro's analysis showed, position is mostly noise; frequency is the signal.
Measuring only one engine. A strong score on ChatGPT tells you nothing about Perplexity or AI Overviews. Measure each separately, or you're optimizing blind.
Counting mentions but ignoring citations (or vice versa). Mentions build association; citations drive traffic. On ChatGPT most mentions carry no link; Perplexity links almost everything. Track them as two separate metrics.
Ignoring sentiment and accuracy. Being named isn't automatically a win. AI systems routinely repeat outdated pricing or wrong positioning. Always log how you're described, not just whether you appear.
Running each prompt only once. Generative answers are probabilistic: the same question names different brands on different runs. One query is an anecdote, not a measurement.
Expecting clean click attribution. Most AI mentions never produce a trackable click, so analytics will always undercount AI's influence. Treat attribution as directional, and add a "Did you hear about us through an AI assistant?" option to your intake to catch the rest.
Testing once and moving on. Visibility drifts as models update and content ages. Re-run the same prompt set on a fixed cadence so you can see what's actually moving.
From measuring to moving
A spot check tells you where you stand today. Improving, and staying improved, takes continuous measurement across every engine, plus the context a manual run can't give you: which sources are driving your visibility, how you compare to competitors over time, and which prompts you're losing and why.
This is where the afternoon spot check runs out of road, and where 3RD picks it up. It runs the same kind of prompt testing continuously and at scale, across hundreds of prompts spanning ChatGPT, Perplexity, Gemini, Claude and Google AI, so your baseline becomes a live trend instead of a one-off snapshot. But the part a manual run can never give you is the why: 3RD traces which specific sources are winning each prompt, where a competitor is pulling ahead, and which piece of content would close the gap. You stop guessing what to fix and start working a prioritized list.
The bottom line
You can't manage what you can't see, and traditional analytics can't see AI answers. Start with a manual spot check using the prompt matrix above, measure each engine on its own terms, track frequency rather than rank, and watch both mentions and citations. Do that consistently and "how visible are we in AI?" stops being a worry and becomes a number you can move.
Want your baseline without the spreadsheet? Get started with 3RD to see how your brand shows up across ChatGPT, Perplexity, Gemini, Claude and Google AI, and where to improve first.