Where Do AI Citations Come From Across Industries?

Where Do AI Citations Come From Across Industries?

Profound classified a proprietary 11.84-billion-citation corpus across eight models, with useful patterns and a hard causal limit.
Profound classified a proprietary 11.84-billion-citation corpus across eight models, with useful patterns and a hard causal limit.

6 min

Zach Chmael

In This Article

An 11.84-billion-citation analysis found source mix varies by industry and model. Use it as a baseline, not a causal content prescription.

Updated

TL;DR

Where Do AI Citations Come From Across Industries?

In Profound's customer corpus, company-operated websites supplied roughly 57% of AI citations. That is the clearest answer. It is also the easiest number to misuse.

The analysis covers 11.84 billion citations from eight answer engines over three months. Its source mix changes sharply by industry and model. The useful conclusion is that teams need to inspect the evidence supply around their own buyer questions. The unsupported conclusion is that moving budget toward whichever source bucket has the largest bar will cause more citations.

A citation-share benchmark describes what appeared in answers gathered under one system's prompts, customers, models, and classification rules. It does not identify why a source was selected or which intervention would change the next answer.

What counts as an AI citation source?

An AI citation source is the web property behind a visible citation, but the category assigned to that property determines what the benchmark can tell you.

Profound normalized each hostname to its registrable domain, then matched it against a map of 3.02 million domains. The map covered 98.3% of global citation volume.

Each matched domain went into one of three main buckets:

Source bucket

Profound's definition

What the label can hide

Brand

A company-operated corporate site, product site, blog, or documentation property

The focal company's site, a competitor, and an adjacent vendor all share one label

Earned

Third-party editorial coverage, institutional sources, and PR wire

Independent reporting, academic material, and paid distribution can sit together

Social

Social and user-generated platforms such as Reddit and YouTube

A product review, support thread, creator video, and casual comment can share one label

Other or unmapped

The remainder outside those three classes

The source may be new, stale in the map, or difficult to classify

That taxonomy is useful for seeing broad source supply. It is too coarse to prescribe work. A competitor's documentation can tell an answer model something true about a category while still being a poor source for your product claims. A PR wire and a regulator's page both count as earned, yet they carry very different evidence weight.

I checked the methodology before using the headline. The residual other and unmapped share ranges from 2% to 7% by industry. Profound also notes that its map is point-in-time, so a newly created or reclassified domain can inherit a stale label or fall outside the mapped set.

Where do AI citations come from by industry?

Citation sources differ materially by industry in this corpus, with company-operated sites leading in most categories but earned sources dominating several high-scrutiny fields.

Profound reports that brand sites supplied the largest median share in 24 of 29 industries. The spread is wide. Cybersecurity reached a 74% median brand-site share, while government and nonprofit sat at 15.9%.


Median AI citation source mix across eight industries

Source: Profound.

The chart shows why the global 57% brand-site share cannot become a universal target. Pharma's median mix is 34% brand, 59% earned, and 5% social. Software is nearly the reverse at 73% brand, 12% earned, and 13% social.

Fashion sits at 50% brand, 31% earned, and 17% social. Marketing reaches 71% brand, while travel and hospitality is 54% brand, 27% earned, and 15% social.

Those differences are real within Profound's sample. Their causes remain open. Industry content supply, customer selection, prompt design, geography, language, retrieval policy, regulation, and classification can all move the mix. The chart does not isolate any one of them.

How does citation mix change by AI model?

Citation mix changes by model because each product has different retrieval systems, data access, interfaces, and answer behavior, though this study cannot isolate which difference caused each bar.

The public analysis includes eight model surfaces: ChatGPT, Claude, Google AI Mode, Google AI Overviews, Google Gemini, Grok, Microsoft Copilot, and Perplexity. Per-model cells needed at least 100 citations, and each model-industry row needed at least 30 cells.


Median citation source mix for eight AI models

Source: Profound.

ChatGPT had the lowest brand-site share among the tracked models at 47%. Google Gemini had the highest at 69%. ChatGPT's earned share reached 30%, while Google AI Overviews was 17%.

Social also varied. Google AI Overviews drew 15.3% from social sources, and Google AI Mode drew 14.4%. Microsoft Copilot used social for about 1 in 29 citations.

A team should record the model surface instead of blending these shares immediately. A single average can hide that the same source type is common in one product and scarce in another. It can also hide a model update, retrieval change, or prompt mix that moved the result without any change to the company's evidence.

What does the 11.84-billion denominator leave out?

The 11.84-billion denominator leaves out market representativeness, prompt design, classification error, regional mix, citation correctness, and buyer outcomes.

The sample began with 8,061 active Profound categories. The final analysis kept 7,542 categories that had at least 100 citations. Those are Profound categories, not a random draw of companies or buyer questions.

Each category received one industry label from an LLM judge across 29 industries. The article does not publish a validation set, human agreement rate, or classification error estimate. It also does not disclose the category names, customer count, prompt set, failed runs, model versions, or citation totals by category.

The aggregation choice is thoughtful: every category receives equal weight, so high-volume categories cannot dominate the industry median. But equal category weighting does not make the sampled categories representative of an industry. It changes what the median means.

I also inspected both charts in their original article context. Their footnotes preserve the 16 April to 16 July 2026 period, model list, source API, and owner-based industry labels. Those details need to travel with the bars. Without them, a point-in-time customer corpus starts to look like a permanent map of the web.

Can citation mix choose a content channel?

Citation mix can identify where to investigate, but it cannot choose a content channel because the analysis contains no assigned intervention or causal comparison.

Profound reports that within one industry, the earned-share spread between the 75th- and 25th-percentile categories can reach 43 percentage points. That is useful evidence of variation. It does not show that a company below the median should buy PR, publish guest posts, start a YouTube program, or produce more owned pages.

The diagnosis should stay closer to the raw answer:

Observed pattern

What it supports

What to inspect next

Competitor-owned pages supply the decisive facts

Competitor evidence is available to the model

Which facts are missing or ambiguous on the focal company's public sources?

Independent sources dominate the answer

Third-party material is present in this source set

Is the coverage current, accurate, and based on verifiable proof?

Social sources shape the comparison

Community content is part of the answer's evidence supply

Are the cited posts representative, current, and correctly interpreted?

The company is cited but not recommended

Citation presence did not settle fit

Which comparison criteria, product facts, or proof shaped the recommendation?

No relevant company source appears

The source set may have an access, relevance, or authority gap

Can the right page be fetched, retrieved, and distinguished for this question?

A citation is not a mention. A mention is not a recommendation. None of those events is revenue. Channel work should begin only after the team knows which event failed and what evidence the answer actually used.

How can Trovance turn citation mix into a useful diagnosis?

Trovance can turn citation mix into a useful diagnosis by attaching source type to the actual buyer question, raw answer, explanation, comparison, and recommendation instead of treating the industry median as a prescription.

A team can observe how AI systems explain, cite, compare, and recommend its company for a defined question. It can then inspect whether the sources are company-owned, competitor-owned, independent, institutional, community-based, or unknown, and whether the answer used each source accurately.

Observation still does not resolve the gap. Trovance diagnoses the evidence problem behind the result, then helps the team produce and publish the proof-backed asset that should exist. The right action might be a product-page refresh, an honest comparison, clearer documentation, independent validation, or no new content.

A team would use Trovance here to keep the prompt, answer, sources, diagnosis, and production decision in one evidence trail. It does not promise a citation or recommendation.

Use Trovance to trace citation mix back to the evidence decision.


FAQs

Where do AI citations usually come from?

In Profound's 11.84-billion-citation corpus, company-operated sites supplied roughly 57% of citations across the tracked categories and models. That includes any company's corporate site, product pages, blogs, and documentation. It is a customer-corpus median, not a universal estimate for every prompt, market, language, or AI product.

What is a brand-site citation?

Profound defines a brand site as a web property operated by a company. The label includes the focal company's pages, competitor sites, adjacent vendors, corporate blogs, product documentation, and other owned properties. It does not mean the cited page belongs to the company being measured or supports that company positively.

Do ChatGPT and Google cite the same source types?

No. In the reported corpus, brand-site share ranged from 47% for ChatGPT to 69% for Google Gemini. Google AI Overviews and AI Mode also used social sources more often than Microsoft Copilot. These are point-in-time descriptive shares, not fixed product rules or proof of different training data.

Which industries rely most on earned media citations?

Profound reports a 59% median earned-media share for pharma and biotech, compared with 11.4% for SaaS and software. Government and nonprofit also leaned heavily toward earned sources. The study does not isolate whether regulation, content supply, prompts, customer selection, geography, or model retrieval policy caused those differences.

Does a high citation share prove a content strategy works?

No. Citation share records which source classes appeared in observed answers. It does not assign companies to a content intervention or compare a treated group with a control. A high owned, earned, or social share therefore cannot prove that publishing more on that channel will increase citations, recommendations, traffic, pipeline, or revenue.

How should a company benchmark its citation sources?

Start with one buyer question and one model surface. Preserve repeated answers, every citation, the date, and the recommendation outcome. Classify sources with enough detail to act, separating your pages from competitors and independent proof from PR wire. Compare with a market baseline only after the local denominator is visible.

What should a team do after finding a citation gap?

Inspect whether the missing source reflects access, retrieval, stale facts, weak proof, unclear comparison language, or legitimate bad fit. Then choose the smallest justified action: refresh an existing page, publish missing evidence, seek independent validation, clarify the product, or publish nothing. Rerun the same protocol without promising an outcome.


Related Resources

Trace citation mix back to the evidence decision with Trovance

Be the answer.

Built to win the agentic web. Made to improve the human world.

Be the answer.

Built to win the agentic web. Made to improve the human world.

Be the answer.

Built to win the agentic web. Made to improve the human world.