How NetPageTwo measures AI citation rate (full methodology)
Most agencies selling AI search optimization don’t tell you how they measure it. They report numbers like “+340% AI visibility” without saying which engines they tested, which queries they ran, or what counted as a citation.
We publish the methodology. The whole thing. If you want to verify the numbers in our case studies or replicate the test on your own brand, this page tells you exactly how.
What we measure
Two metrics for every client, tracked separately for each of 5 AI engines.
Citation rate is the percentage of test queries where the client’s brand appears in the AI response. If we run 50 queries and the brand is mentioned in 12 of them, citation rate is 24%. Binary at the query level: either the brand shows up or it doesn’t.
Citation prominence is how prominently the brand appears when it does. Five-point scale: 0 (not present), 1 (supporting reference, listed but not named in synthesis), 2 (named in synthesis but secondary), 3 (primary citation, named in first paragraph), 4 (authoritative source, the synthesis is built around this brand). We compute the average across queries where the brand appeared.
The two numbers describe the AI search position. Citation rate tells you how often. Citation prominence tells you how well.
Which engines we test
Five. In priority order for B2B clients:
- ChatGPT browse mode (OpenAI). Largest user base. Generally the highest-impact channel for B2B.
- Perplexity. Smaller user base but high-intent queries.
- Google AI Overviews. The AI summary at the top of regular Google search.
- Gemini (Google). Integrated into Workspace.
- Microsoft Copilot. Enterprise-skewed buyer base.
For clients with specific buyer demographics, we also test:
– Grok (xAI): tech, crypto, X-heavy operator audiences
– DeepSeek, developer-heavy and international audiences
– Meta AI, consumer brands and local services
We track each engine separately because they cite differently. Aggregating loses signal.
How we pick test queries
For each client, we build a 50-query test set covering five categories.
Commercial intent (15 queries): “best [category] for [use case],” “what is the best [category],” “top [category] companies in 2026.” Direct purchase consideration queries.
Comparison intent (10 queries): “[client] vs [competitor],” “[competitor] alternative,” “is [client] better than [competitor].” Direct comparison queries.
Problem-aware (10 queries): “how to solve [problem the client addresses],” “what to do when [pain point].” Pre-purchase research queries.
Brand-specific (10 queries): “what does [client] do,” “is [client] worth it,” “does [client] offer [feature].” Direct brand investigation.
Solution-aware (5 queries): “what features should I look for in [category],” “what does [category] cost in 2026.” Late research stage.
The exact query phrasing comes from real prospect conversations, sales call recordings, and customer support tickets when available. We don’t invent queries from generic templates.
How we score
Each query gets run through each engine. We capture:
- The full AI response text
- The list of cited sources (if the engine surfaces them)
- Where the brand appears in the response (first paragraph, middle, end, source list only, or not at all)
- Whether the response cites the client specifically vs cites a third party who mentions the client
Scoring is two-pass.
Pass 1: Citation rate. Each query gets a 0 or 1. 1 if the brand is mentioned anywhere in the AI response. 0 if not. The brand’s product name, founder name, or branded URL all count.
Pass 2: Citation prominence. For queries that scored 1, we apply the 0-4 prominence scale. Manual review for each. We don’t use automated NLP for prominence scoring because the subjective signal of “this brand is named as the authority” is too important to delegate to a model.
The full scoring takes 2-3 hours per audit for 50 queries × 5 engines = 250 query-engine combinations. We do it by hand for accuracy.
What counts as a citation (and what doesn’t)
Three things count:
- The brand name appears in the response text
- A URL from the brand’s domain is cited in the source list
- A direct product name owned by the brand appears in the response
Three things don’t count:
- The brand appears only via a third-party article (e.g., “Forbes mentioned [brand]”): that’s a Forbes citation, not a brand citation
- The brand is mentioned in passing without being the answer to the query
- The brand is mentioned only in a disclaimer or footnote
We separately track third-party citations because they matter for SEO and entity authority, but they’re not the primary metric.
Sample query set (for a hypothetical B2B SaaS CRM client)
To make this concrete, here’s a sample of the 50 queries we’d run for a B2B SaaS CRM company:
- “best B2B CRM for small business 2026”
- “Salesforce alternative for SMB”
- “what is the cheapest B2B CRM”
- “HubSpot vs Pipedrive 2026”
- “does [client] integrate with QuickBooks”
- “how to choose a CRM for a 20-person team”
- “what features should I look for in a B2B CRM”
- “is [client] worth $200 a month”
- “[client] pricing tiers explained”
- “best CRM for outbound sales 2026”
The full 50-query set gets shared with the client at the start of the engagement so they can sanity-check the query selection.
What the report looks like
Monthly. One spreadsheet plus a 5-minute Loom walkthrough. Columns include:
- Query text
- Engine
- Citation rate (0/1)
- Prominence score (0-4)
- Source URL cited (if any)
- Response excerpt (the part mentioning the brand)
- Trend vs previous month
We also produce an aggregate dashboard showing:
– Total citation rate by engine
– Citation prominence average by engine
– Top citation surfaces (which pages of the client’s site are getting cited)
– Top wasted opportunities (queries where the brand should appear but doesn’t, with diagnosis)
The client gets full access to the raw data so they can audit our scoring.
What we don’t claim
Three honest limits.
We don’t claim our scoring is the industry standard. There isn’t one yet for AI citation measurement. Other agencies use different methodologies. Compare ours to theirs before deciding which is more defensible for your situation.
We don’t claim 50 queries is comprehensive. It’s a sample. The real query universe for a typical B2B brand is in the thousands. We optimize the 50 to cover the highest-intent, highest-volume queries.
We don’t claim AI citation rate predicts revenue with precision. It correlates with inbound but the conversion path from “cited in ChatGPT” to “qualified opportunity” varies by buyer demographic and product. We track citation as an input metric, not a revenue metric.
Why we publish this
Two reasons.
If our methodology can be audited, our results are defensible. The AI search optimization vertical has plenty of vendors making claims that fall apart under scrutiny. We’d rather publish the test design and let clients verify than rely on opacity.
If the methodology is good, it’ll get cited by AI engines and by third parties who reference how to measure AI citation. That’s both honest distribution and indirect entity authority for the brand.
If you want to run this test on your own domain before working with us, the query selection logic above is enough to replicate it. We’ll publish the scoring rubric as a downloadable template separately.
If you want us to run the test and deliver the audit, the fit call covers it. The audit ships in week one with full methodology disclosed.
Sources and further reading
- Google Search Central: AI Overviews: Google’s official documentation on how AI Overviews surface citations.
- Perplexity citations methodology: how Perplexity selects sources for its synthesized answers.
- The Princeton GEO research paper: the foundational academic work on Generative Engine Optimization measurement.
Related reading:
– AI search visibility vs ranking: the distinction that changes the work
– How ChatGPT actually decides what to cite (three signals tested)
– Case studies (real numbers, verified in client GA4)
– AI engines hub
– Industries
GEO 101