How to measure AI search visibility: prompt sets, mention rate, citation share
Shivam GuptaPublished Updated 8 min
Build a fixed set of buyer-intent prompts for your market, run it on a schedule across ChatGPT, Gemini, Perplexity, Claude, Google AI Overviews and Meta AI, and log two numbers per engine: mention rate (how often you're named) and citation share (your share of cited sources versus competitors). Track the trend monthly against a baseline.
Why can't you just check your analytics?
Because most of AI search's effect never touches your analytics. An engine can recommend you to a buyer who calls without clicking, or recommend your competitor to a buyer you never knew existed. Referral traffic from AI domains is real but structurally undercounts influence — it's a floor, not a measure. Measurement has to happen in the answers themselves.
There's also no vendor console coming to save you: no engine publishes a Search Console equivalent showing your impressions in its answers. That absence is why we built a measurement routine and treat it as the spine of every engagement — and it's the honest reason this article exists: what follows is our actual methodology, publishable because a methodology you have to hide isn't one. The zero-click dynamics make this more pressing every year the click share falls.
What is a prompt set and how do you build one?
A prompt set is a fixed basket of 25–50 questions your buyers actually ask, phrased the way they ask them — problem-first, constraints included. It spans the funnel: problem questions ('should I get a lawyer for a car accident that wasn't my fault'), selection questions ('best CRM for a 10-person sales team'), and branded prompts about you and named competitors.
Two construction rules carry most of the value. Source prompts from reality — sales-call questions, support tickets, autocomplete patterns, community threads — not from your keyword list, because buyers prompt engines in sentences keywords never predicted. And fix the basket: the set must survive unchanged for quarters, because every prompt you swap breaks the trend line at exactly the moment you most want to trust it. Vendors who rotate prompts mid-engagement are, deliberately or not, making their results uncomparable to their baseline.
Size the set to what you can sustain, not to what sounds thorough. Twenty-five prompts run reliably every month beat eighty prompts run once — and every prompt added multiplies runs across six engines and repeated samples. Start narrow on the questions closest to revenue; widen the basket only when the routine is boringly stable.
What is mention rate?
Mention rate is the percentage of answers across your prompt set in which the engine names your business — linked or not. Forty prompts, fourteen answers naming you: 35% on that engine, that run. It's the closest thing AI search has to a rank position, and it's the number that makes 'are we visible' answerable.
Named-not-linked is counted deliberately: an unlinked recommendation still steers the buyer, invisibly to analytics. Log the answer text itself too — how you're characterized matters, and for regulated businesses the verbatim record is worth keeping in its own right.
What is citation share?
Citation share is your domain's percentage of all sources cited across your prompt set, per engine, against competitors. Where mention rate measures whether you're in the recommendation, citation share measures whether your pages are the evidence behind answers. It's the leading indicator: engines typically cite a domain before they start naming the business unprompted.
The two metrics diverge in diagnostic ways. Cited-but-not-mentioned suggests quotable content attached to a weak brand entity; mentioned-but-not-cited suggests your reputation lives on third-party sources while your own pages aren't worth quoting. Each pattern prescribes different work — which is precisely what a measurement system is for.
Why measure every engine separately?
Because engines agree on almost nothing. Only 11% of domains cited by ChatGPT are also cited by Perplexity (AuthorityTech). A blended score averages away the information you need: which engine you've won, which you're absent from, and where the next month of work should aim.
11%
of domains cited by ChatGPT are also cited by Perplexity
Each engine has distinct citation logic — per-engine tracking required
Source: AuthorityTechPer-engine data also catches surface-specific behavior — an on-site fix landing on Perplexity weeks before ChatGPT reflects it, or Google AI Overviews moving with your classic index health while chat engines lag. We track six engines (ChatGPT, Gemini, Perplexity, Claude, Google AI Overviews, Meta AI) because that's where the answer volume is; the right roster for you depends on where your buyers actually ask.
What does a monthly report look like?
One page per month, honestly filled: mention rate and citation share per engine with deltas against baseline, the specific prompts that moved, the prompts that didn't, and the work shipped that month so movement can be attributed rather than implied. Negative results stay in the report — a month where nothing moved is information, not failure to be edited out.
Run cadence: monthly is the honest floor for trend reading; because engine outputs vary run to run, each cycle uses repeated runs rather than single samples, so deltas reflect movement rather than noise. Anyone can operate this methodology with patience and a spreadsheet — genuinely — and the monitoring service exists for businesses that want the routine plus the prioritized fixes it points to, delivered as one motion.
Related questions, answered straight
What's a good mention rate?
There's no universal benchmark, and anyone quoting one is selling something. Meaningful comparisons are your own baseline (are you trending up?), your named competitors on identical prompts (who owns your market's answers?), and per-engine spread (visible on Perplexity but absent from ChatGPT points to specific work). A 20% mention rate can be dominant in one market and weak in another.
How often should the prompt set run?
Monthly for trend reporting, with repeated runs per cycle rather than single samples — engine outputs vary run to run, and repetition separates movement from noise. Weekly runs make sense mid-sprint or when watching a specific fix land. Rarer than monthly and you can't attribute changes to the work that caused them, which defeats the purpose of measuring.
Can't a tool like Profound or Otterly do this measurement for me?
Yes — running prompts and dashboarding mentions is exactly what they're built for, from around $25–99/mo, and if you have someone to act on the findings, that's a fine setup. What tools don't provide is the acting: restructuring, schema, entity cleanup, earned citations. We run the same kind of tracking and then do the work it points to.