Why Prompt Tracking Is the Wrong Metric
Most teams measure AI search visibility by counting how often their brand appears in AI-generated answers. This is the wrong number. Prompt tracking—typing a set of prompts into ChatGPT, Perplexity, or Google AI Overviews—feels like rank tracking, but it isn't. It measures citations, not recommendations. A citation is a footnote; a recommendation is a directive to choose your business. The gap between them is where your revenue lives.
Citations vs. Recommendations: The Critical Distinction
Lily Ray's analysis of 100 business software queries found that when a brand's own listicle was cited, the brand was left out of the actual recommendation 69% of the time. Jeff Oxford's team tested 20,000 ChatGPT responses and found only a 0.4 correlation between being cited and being recommended. BrightEdge reported source overlap between engines as low as 16%, but recommended brands stayed in a tighter 36–55% band. Being a source does not mean being chosen.
The Instability of AI Answers
Rand Fishkin's SparkToro study quantified the problem: to get two identical lists of brands in an AI answer, you'd need to ask Claude or ChatGPT 1,500 times. Single-shot measurements are noise. AI answers vary by prompt, user, and model version. Reliable measurement requires statistical sampling—like polling, not rank checking.
What to Measure Instead: Presence and Recommendation Share
Replace prompt tracking with presence: how often your brand is named across the answer space, and whether that presence converts into a recommendation. Fishkin calls this 'percent of visibility'—like brand awareness surveys. Wil Reynolds adds that you must also track answer composition (e.g., word count) because when models double answer length, raw visibility inflates without real value. The ultimate metric is action: visibility only matters if it drives clicks or conversions.
Start with Brand Accuracy
Before measuring recommendations, audit whether AI describes your entity correctly. Alisa Scharf's brand accuracy audit scores models on objective facts (founding date, location, products). If the model holds wrong facts, every downstream metric is built on sand. Duane Forrester advises becoming the canonical source—the trusted answer—so the machine has no reason to switch.
Blind Spots: Training Cutoff and Platform Data
Two blind spots remain: training-data cutoff (models may rely on stale knowledge) and lack of platform data from OpenAI or Anthropic. Google's Search Console AI impressions report is weak but available. Microsoft's Bing Webmaster Tools offers similar data. The frontier model companies have little incentive to share usage data, leaving a gap in measurement.
Your Move: Audit, Align, and Measure
Start with a brand accuracy audit. Ensure consistent entity descriptions across your website, schema, social profiles, and third-party mentions. Then measure presence and recommendation share using statistical sampling (multiple prompts, multiple times). Track actions (clicks, conversions) tied to AI visibility. Finally, monitor legal developments: a German court held Google liable for false AI statements, suggesting platforms will prioritize confident, accurate entities.
FAQ
It counts citations, not recommendations. A citation is a footnote; a recommendation drives action. Most tools conflate the two.
Use statistical sampling—multiple prompts, multiple times—to measure presence (how often you're named) and recommendation share (how often you're recommended). Track actions like clicks.
Run a brand accuracy audit: check if AI describes your entity correctly (founding date, location, products). Fix inaccuracies before optimizing for recommendations.


