AI search products can cite manufacturers, retailers, marketplaces, directories, review publications, forums, primary research and other sources for buying-intent questions. Which source type appears depends on the product, prompt, market, time and available evidence.
A defensible study must sample those conditions and classify citations consistently. One set of branded prompts cannot support a universal “most cited websites” ranking.
Define the observation unit before collecting answers
Use one row per product × mode × prompt × market × language × scheduled run. This original SEO Companies Hub schema keeps the denominator visible.
| Field group | Required fields | Why it matters |
|---|---|---|
| Product context | Product, mode, plan/account state, interface and collection time | Prevents unlike retrieval experiences being pooled |
| Prompt context | Frozen prompt ID, buying stage, category, buyer constraints and brand status | Makes coverage and prompt bias auditable |
| Market context | Country, language, currency and known location state | Preserves price, availability and legal scope |
| Citation event | Citation type, position, URL, page role and supported claim | Separates inline support from source lists and brand mentions |
| Review result | Accessible, claim present, representation accurate, current and limitations preserved | Prevents raw citation count from rewarding misuse |
| Outcome context | Recommendation, referral if observable and customer outcome if lawfully joinable | Separates visibility from business value and causality |
A leaderboard without these fields is a convenience sample, not a market-wide fact.
Define buying intent
Use a staged taxonomy:
Category exploration
“What type of accounting software works for a small nonprofit?”
Criteria research
“What should I compare when choosing an SEO agency?”
Product or provider comparison
“Compare product A and product B for a ten-person team.”
Price and availability
“How much does this service cost in Kenya?”
Validation
“Is this provider suitable for a regulated business?”
Transaction readiness
“Where can I buy or book this today?”
The source needs change by stage. Primary product documentation may dominate specifications, while independent reviews and directories support comparisons.
Define the research products
Name every product and mode:
- ChatGPT with web search;
- Perplexity mode or model;
- Google AI Overviews or AI Mode;
- another assistant or connected source;
- desktop, mobile or browser experience.
Record plan, sign-in state and relevant settings where possible. Product behavior can vary and change.
OpenAI's ChatGPT search documentation describes web search and cited responses. Perplexity documents distinct crawlers and user-triggered fetch behavior, while Google documents AI features within Search. These are different retrieval systems; do not combine results without product labels.
Build a prompt frame
Create variables:
- product or service category;
- buyer role;
- company or household size;
- geography;
- budget;
- use case;
- integration or constraint;
- decision stage;
- brand-known or brand-neutral wording;
- freshness need.
Use a balanced sample across categories and stages. Avoid prompts constructed around the publisher's own brand unless branded behavior is a separate research question.
Freeze the prompt set before collecting results. Version changes rather than replacing inconvenient prompts.
Sample markets and languages
Buying answers are local. Price, inventory, law and provider availability vary.
Record:
- country and city where relevant;
- language;
- currency;
- date and local time;
- device or interface;
- location personalization known;
- account state;
- product availability.
Do not generalize United States English results worldwide. Use qualified local reviewers for interpretation.
Repeat observations
Generated responses can vary. Run each prompt more than once across scheduled windows, subject to product terms and research ethics.
For every response, capture:
- full prompt;
- answer;
- inline citations;
- source list or related links;
- cited URL and domain;
- quoted or supported claim;
- answer recommendation;
- errors or refusal;
- collection timestamp.
Preserve raw evidence and a hash or stable ID. Do not collect personal account data unnecessarily.
Define what counts as a citation
Separate:
- inline citation supporting a claim;
- source-list link;
- related result;
- image source;
- unlinked brand mention;
- product or app card;
- navigational link;
- advertiser placement if present.
Count each according to a declared rule. A domain appearing in “related links” is not the same as an inline source.
Deduplicate repeated links within one response while retaining citation frequency and positions.
Create a source taxonomy
Classify at page level:
- manufacturer or official provider;
- retailer or reseller;
- marketplace;
- directory or aggregator;
- independent review publication;
- news or specialist media;
- government or regulator;
- standards body;
- academic or research institution;
- community forum;
- social platform;
- personal expert site;
- affiliate publisher;
- pricing or data platform;
- unknown.
Domains can serve multiple roles. Classify the cited page's function for that claim, not only the company brand.
Document commercial relationships when observable.
Evaluate citation support
Generated citations can be incomplete, outdated or incorrect. The OpenAI search documentation tells users to open and review cited sources; every citation in this study therefore needs a manual support check.
Rate:
- source accessible;
- claim present;
- claim accurately represented;
- date suitable;
- geographic scope correct;
- primary versus secondary;
- methodology available;
- material limitation preserved;
- price or availability current;
- commercial disclosure visible.
A citation count without accuracy rewards sources even when the answer misuses them.
Separate citation and recommendation
A source can support a factual claim while the answer recommends a different product. Record:
- cited brand;
- recommended brand;
- comparison criteria;
- positive, neutral or negative context;
- whether the source is the seller;
- whether the answer discloses uncertainty.
Do not label a citation as endorsement. A competitor can be cited as evidence for its own price in an unfavorable comparison.
Measure at several levels
Domain citation share
Percentage of eligible responses containing the domain under the stated citation rule.
Page citation frequency
Unique cited URLs and repeated use.
Source-type share
Distribution among official, review, marketplace, government and other categories.
Accuracy rate
Supported citations divided by reviewed citations.
Prompt coverage
Citation presence by buying stage, market and category.
Recommendation association
How often a cited source is linked to a recommendation, without assuming causation.
Show counts and confidence. Responses are nested within prompts and products; account for that structure.
Avoid a biased leaderboard
Rankings can be distorted by:
- more prompts in one category;
- repeated domains serving many products;
- branded prompts;
- one market;
- one day;
- one product mode;
- source-list links counted as citations;
- duplicate URLs;
- unreachable sources;
- prompt wording that names likely sources.
Weight categories transparently or publish stratified tables. Do not invent a universal score from a convenience sample.
Analyze why page types are useful
Inspect cited pages for observable features:
- primary product facts;
- current price and availability;
- consistent comparison criteria;
- first-hand testing evidence;
- original data;
- clear author and organization;
- methodology;
- dates and updates;
- accessible page structure;
- stable URL;
- independent corroboration.
This analysis creates hypotheses, not hidden ranking factors. A feature can be common among cited pages without causing citation.
Improve publisher citation readiness
For buying-intent pages:
- expose product and service scope;
- publish accurate price or cost drivers;
- keep availability and location current;
- use comparable specifications;
- show test method and limitations;
- distinguish seller claims and independent evidence;
- use accessible tables and text;
- maintain stable canonical URLs;
- link to primary sources;
- disclose affiliate and commercial relationships;
- provide useful next actions.
Do not rewrite the page to mimic generated answer fragments. Serve the buyer's decision.
Measure business impact separately
Track referrals from identified products, landing pages, task completion, valid leads, transactions and net value. Connect citations from the research sample only when the corresponding link and period can be observed.
A highly cited directory may send few visits. A rarely cited product page may send high-intent customers. Report both visibility and customer value.
Do not claim a citation caused a sale without evidence.
Publish the method and data rights
A report should include:
- products and modes;
- prompt inventory;
- markets and dates;
- collection procedure;
- citation definition;
- source taxonomy;
- support-review rubric;
- deduplication;
- weighting and statistics;
- sample and missing responses;
- limitations;
- data and screenshot rights;
- correction policy.
Protect user and account data. Follow product terms for automated collection.
A practical verdict
AI products cite many source types for buying questions, and the mix changes with intent and context. The correct research task is not to declare one permanent winner but to measure which sources appear, whether they support the claims and how patterns differ by product, market and buying stage.
Publish the raw method and limitations. Use findings to improve primary evidence and buyer usefulness, not to manufacture citations.
Related decisions
- GEO Strategy: How to Build Evidence AI Search Systems Can Cite — the adjacent ai search decision.
- Agentic Search Optimization Explained: From Discovery to Selection — the adjacent ai search decision.
- Does AI-Generated Content Work for SEO? An Evidence and QA Framework — the adjacent ai search decision.