
6 min
Zach Chmael
In This Article
An 11.84-billion-citation analysis found source mix varies by industry and model. Use it as a baseline, not a causal content prescription.
Updated
TL;DR
🧾 Profound analyzed 11.84 billion citations gathered from 16 April through 16 July 2026.
🧭 The final analysis covered 7,542 categories with at least 100 citations.
🏢 Brand-operated sites supplied roughly 57% of citations across the tracked corpus.
🔀 Brand-site share ranged from 47% for ChatGPT to 69% for Google Gemini.
🧪 The domain map classified 3.02 million domains and 98.3% of citation volume, but the study was observational.
Where Do AI Citations Come From Across Industries?
In Profound's customer corpus, company-operated websites supplied roughly 57% of AI citations. That is the clearest answer. It is also the easiest number to misuse.
The analysis covers 11.84 billion citations from eight answer engines over three months. Its source mix changes sharply by industry and model. The useful conclusion is that teams need to inspect the evidence supply around their own buyer questions. The unsupported conclusion is that moving budget toward whichever source bucket has the largest bar will cause more citations.
A citation-share benchmark describes what appeared in answers gathered under one system's prompts, customers, models, and classification rules. It does not identify why a source was selected or which intervention would change the next answer.
What counts as an AI citation source?
An AI citation source is the web property behind a visible citation, but the category assigned to that property determines what the benchmark can tell you.
Profound normalized each hostname to its registrable domain, then matched it against a map of 3.02 million domains. The map covered 98.3% of global citation volume.
Each matched domain went into one of three main buckets:
Source bucket | Profound's definition | What the label can hide |
|---|---|---|
Brand | A company-operated corporate site, product site, blog, or documentation property | The focal company's site, a competitor, and an adjacent vendor all share one label |
Earned | Third-party editorial coverage, institutional sources, and PR wire | Independent reporting, academic material, and paid distribution can sit together |
Social | Social and user-generated platforms such as Reddit and YouTube | A product review, support thread, creator video, and casual comment can share one label |
Other or unmapped | The remainder outside those three classes | The source may be new, stale in the map, or difficult to classify |
That taxonomy is useful for seeing broad source supply. It is too coarse to prescribe work. A competitor's documentation can tell an answer model something true about a category while still being a poor source for your product claims. A PR wire and a regulator's page both count as earned, yet they carry very different evidence weight.
I checked the methodology before using the headline. The residual other and unmapped share ranges from 2% to 7% by industry. Profound also notes that its map is point-in-time, so a newly created or reclassified domain can inherit a stale label or fall outside the mapped set.
Where do AI citations come from by industry?
Citation sources differ materially by industry in this corpus, with company-operated sites leading in most categories but earned sources dominating several high-scrutiny fields.
Profound reports that brand sites supplied the largest median share in 24 of 29 industries. The spread is wide. Cybersecurity reached a 74% median brand-site share, while government and nonprofit sat at 15.9%.

Source: Profound.
The chart shows why the global 57% brand-site share cannot become a universal target. Pharma's median mix is 34% brand, 59% earned, and 5% social. Software is nearly the reverse at 73% brand, 12% earned, and 13% social.
Fashion sits at 50% brand, 31% earned, and 17% social. Marketing reaches 71% brand, while travel and hospitality is 54% brand, 27% earned, and 15% social.
Those differences are real within Profound's sample. Their causes remain open. Industry content supply, customer selection, prompt design, geography, language, retrieval policy, regulation, and classification can all move the mix. The chart does not isolate any one of them.
How does citation mix change by AI model?
Citation mix changes by model because each product has different retrieval systems, data access, interfaces, and answer behavior, though this study cannot isolate which difference caused each bar.
The public analysis includes eight model surfaces: ChatGPT, Claude, Google AI Mode, Google AI Overviews, Google Gemini, Grok, Microsoft Copilot, and Perplexity. Per-model cells needed at least 100 citations, and each model-industry row needed at least 30 cells.

Source: Profound.
ChatGPT had the lowest brand-site share among the tracked models at 47%. Google Gemini had the highest at 69%. ChatGPT's earned share reached 30%, while Google AI Overviews was 17%.
Social also varied. Google AI Overviews drew 15.3% from social sources, and Google AI Mode drew 14.4%. Microsoft Copilot used social for about 1 in 29 citations.
A team should record the model surface instead of blending these shares immediately. A single average can hide that the same source type is common in one product and scarce in another. It can also hide a model update, retrieval change, or prompt mix that moved the result without any change to the company's evidence.
What does the 11.84-billion denominator leave out?
The 11.84-billion denominator leaves out market representativeness, prompt design, classification error, regional mix, citation correctness, and buyer outcomes.
The sample began with 8,061 active Profound categories. The final analysis kept 7,542 categories that had at least 100 citations. Those are Profound categories, not a random draw of companies or buyer questions.
Each category received one industry label from an LLM judge across 29 industries. The article does not publish a validation set, human agreement rate, or classification error estimate. It also does not disclose the category names, customer count, prompt set, failed runs, model versions, or citation totals by category.
The aggregation choice is thoughtful: every category receives equal weight, so high-volume categories cannot dominate the industry median. But equal category weighting does not make the sampled categories representative of an industry. It changes what the median means.
I also inspected both charts in their original article context. Their footnotes preserve the 16 April to 16 July 2026 period, model list, source API, and owner-based industry labels. Those details need to travel with the bars. Without them, a point-in-time customer corpus starts to look like a permanent map of the web.
Can citation mix choose a content channel?
Citation mix can identify where to investigate, but it cannot choose a content channel because the analysis contains no assigned intervention or causal comparison.
Profound reports that within one industry, the earned-share spread between the 75th- and 25th-percentile categories can reach 43 percentage points. That is useful evidence of variation. It does not show that a company below the median should buy PR, publish guest posts, start a YouTube program, or produce more owned pages.
The diagnosis should stay closer to the raw answer:
Observed pattern | What it supports | What to inspect next |
|---|---|---|
Competitor-owned pages supply the decisive facts | Competitor evidence is available to the model | Which facts are missing or ambiguous on the focal company's public sources? |
Independent sources dominate the answer | Third-party material is present in this source set | Is the coverage current, accurate, and based on verifiable proof? |
Social sources shape the comparison | Community content is part of the answer's evidence supply | Are the cited posts representative, current, and correctly interpreted? |
The company is cited but not recommended | Citation presence did not settle fit | Which comparison criteria, product facts, or proof shaped the recommendation? |
No relevant company source appears | The source set may have an access, relevance, or authority gap | Can the right page be fetched, retrieved, and distinguished for this question? |
A citation is not a mention. A mention is not a recommendation. None of those events is revenue. Channel work should begin only after the team knows which event failed and what evidence the answer actually used.

How can Trovance turn citation mix into a useful diagnosis?
Trovance can turn citation mix into a useful diagnosis by attaching source type to the actual buyer question, raw answer, explanation, comparison, and recommendation instead of treating the industry median as a prescription.
A team can observe how AI systems explain, cite, compare, and recommend its company for a defined question. It can then inspect whether the sources are company-owned, competitor-owned, independent, institutional, community-based, or unknown, and whether the answer used each source accurately.
Observation still does not resolve the gap. Trovance diagnoses the evidence problem behind the result, then helps the team produce and publish the proof-backed asset that should exist. The right action might be a product-page refresh, an honest comparison, clearer documentation, independent validation, or no new content.
A team would use Trovance here to keep the prompt, answer, sources, diagnosis, and production decision in one evidence trail. It does not promise a citation or recommendation.
Use Trovance to trace citation mix back to the evidence decision.
FAQs
Where do AI citations usually come from?
In Profound's 11.84-billion-citation corpus, company-operated sites supplied roughly 57% of citations across the tracked categories and models. That includes any company's corporate site, product pages, blogs, and documentation. It is a customer-corpus median, not a universal estimate for every prompt, market, language, or AI product.
What is a brand-site citation?
Profound defines a brand site as a web property operated by a company. The label includes the focal company's pages, competitor sites, adjacent vendors, corporate blogs, product documentation, and other owned properties. It does not mean the cited page belongs to the company being measured or supports that company positively.
Do ChatGPT and Google cite the same source types?
No. In the reported corpus, brand-site share ranged from 47% for ChatGPT to 69% for Google Gemini. Google AI Overviews and AI Mode also used social sources more often than Microsoft Copilot. These are point-in-time descriptive shares, not fixed product rules or proof of different training data.
Which industries rely most on earned media citations?
Profound reports a 59% median earned-media share for pharma and biotech, compared with 11.4% for SaaS and software. Government and nonprofit also leaned heavily toward earned sources. The study does not isolate whether regulation, content supply, prompts, customer selection, geography, or model retrieval policy caused those differences.
Does a high citation share prove a content strategy works?
No. Citation share records which source classes appeared in observed answers. It does not assign companies to a content intervention or compare a treated group with a control. A high owned, earned, or social share therefore cannot prove that publishing more on that channel will increase citations, recommendations, traffic, pipeline, or revenue.
How should a company benchmark its citation sources?
Start with one buyer question and one model surface. Preserve repeated answers, every citation, the date, and the recommendation outcome. Classify sources with enough detail to act, separating your pages from competitors and independent proof from PR wire. Compare with a market baseline only after the local denominator is visible.
What should a team do after finding a citation gap?
Inspect whether the missing source reflects access, retrieval, stale facts, weak proof, unclear comparison language, or legitimate bad fit. Then choose the smallest justified action: refresh an existing page, publish missing evidence, seek independent validation, clarify the product, or publish nothing. Rerun the same protocol without promising an outcome.
Related Resources
Trace citation mix back to the evidence decision with Trovance

