
8 min
Zach Chmael
In This Article
A relevant page can still omit one condition that determines the answer. Use this evidence-gap audit before creating or refreshing another asset.
Updated
TL;DR
๐ The study tested 1,000 multi-condition queries across long documents with evidence on separate pages.
๐ Its corpus contained 2,021 documents and 41,441 rendered pages.
๐ฏ The strongest displayed hybrid found a complete document for 81.1% of queries but put it before partial matches for only 35.8%.
๐งฉ Condition-wise retrieval improved two dense systems by 6.8 and 7.3 points.
๐ Two page-aware systems surfaced all stored support on only 5.1% to 5.3% of queries.
Relevant Content Can Still Miss The Evidence AI Needs
A retriever can find the right document and still fail the buyer's full question. In a July 2026 arXiv preprint, the strongest displayed system found an all-condition document for 81.1% of 1,000 queries, yet ranked that complete document ahead of known partial matches only 35.8% of the time.
That 45.3-point gap is the useful finding. Discovery, coverage, and evidence delivery are separate events. A page can look relevant because it covers the category, product, or use case while omitting one constraint that decides whether the company belongs in the answer.
For a marketing team, the next move is not to merge everything into one enormous page. It is to identify the weakest unsupported condition, find out whether current proof exists and is accessible, then choose the smallest justified action: refresh, consolidate, create, or leave the content alone.
What did the cross-page retrieval study actually test?
The study tested whether retrieval systems can rank a long document containing every requested condition ahead of natural documents that contain only a subset.
Sungguk Cha and five LG Uplus coauthors built CrossPage from 1,753 World Bank PDFs and 268 PMC Open Access articles. The benchmark contains 1,000 requests, including 814 with two conditions and 186 with three. An example asks for a document containing an executive summary, COVID-19, and sanctions.
Every complete document has support for each condition on a different page. The benchmark contains 2,176 all-condition gold judgments and 48,819 subset judgments. The shortest available gold has a median length of 67 pages, so this is a deliberately hard long-document setting.
The primary metric is strict. A result succeeds only when an all-condition document appears in the top 10 and precedes every released partial match. The paper calls this n-Clue Score@10. Gold Hit@10 asks the easier question: did the system find any complete document?
I downloaded the 30-page paper and supplement because the abstract compresses those two events into one paragraph. The tables make the distinction cleaner. A system can score a hit while still putting the incomplete option first.

Source: Do Current Retrievers Cover All The Evidence?.
The figure captures the failure without retrieval jargon. The gold document has A, B, and C. The subset has A and B. If the subset ranks first, finding the gold at rank two does not complete the request.
How large was the gap between finding evidence and covering the request?
The gap remained large across every tested approach, including lexical, dense, visual, decomposed, fused, and reranked systems.
On the clean-query scorecard, BM25-AND found a gold for 57.2% of queries and completed the ranking event for 26.8%. A page-aware ColQwen variant reached 71.2% Gold Hit and 25.2% complete-first success.
The strongest displayed lexical-visual hybrid reached 81.1% Gold Hit and 35.8% n-Clue Score. Its outcome breakdown is even more useful: 18.9% of queries had no gold in the top 10, while 45.3% found a gold but ranked a subset first.
That means candidate discovery was no longer the largest failure. Ordering complete evidence above a fluent, close, incomplete result was.

Source: Do Current Retrievers Cover All The Evidence?.
The controlled contrasts point in the same direction. Splitting the request into condition-specific retrieval lists improved Jina by 6.8 points, with a corrected p-value of .005. It improved Qwen by 7.3 points, with a corrected p-value of .002. Lexical-visual fusion added 8.7 points.
Simply finding more candidates did less. Distinct-page BM25 added 27.9 Gold Hit points but only 1.5 n-Clue points. Scaling one Qwen embedding family from 0.6B to 8B changed complete-first success by 0.0 points, with a 95% interval from -2.7 to 2.8.
Those are bounded results from one benchmark. They do not prove that larger models cannot help or that one retrieval architecture wins everywhere. The SourceShift test is a useful brake: dense decomposition and fusion kept their direction, while page aggregation and lexical composition flipped.
Why does this matter for a company's content system?
It matters because many buyer questions are conjunctions, even when they look like ordinary category queries.
"Which analytics platform supports EU data residency, audit-ready access controls, and our warehouse?" is not three unrelated prompts. A company belongs in the answer only if all three conditions are true, public, current, and supported.
A category page may establish topical relevance. A security page may prove the control. Documentation may list the integration.
A regional hosting note may carry the residency boundary. The proof exists, but the answer system still has to retrieve it, connect it to the same company and offer, and avoid a partial competitor that covers two conditions more clearly.
The study does not test websites or answer engines. Its evidence lives across pages within long PDFs. That boundary matters. Still, its measurement model gives content teams a better audit than "does a relevant page exist?"
Use four states for each buyer condition:
Condition state | What the team can observe | Production decision |
|---|---|---|
Supported and surfaced | The answer states the fact and points to current proof | Preserve and monitor |
Supported but absent | Public proof exists, but the observed answer omits it | Improve access, connection, or placement before creating more prose |
Claimed but unsupported | The answer or site asserts the fact without inspectable proof | Build or qualify the evidence |
Not true or bad fit | The condition is unsupported because the offer does not meet it | Do not manufacture content to force inclusion |
This is the part dashboards tend to skip. Presence can hide an incomplete reason. Absence can hide proof that already exists. Neither result tells a team what to produce until the conditions are inspected separately.
What should a team produce after an incomplete answer?
The team should produce only the smallest asset justified by the unsupported condition and the proof already available.
Start by separating an evidence gap from an access, positioning, product, or fit problem. If current proof already exists, the work may be a documentation repair or a clearer connection between pages. If the claim is true but unsupported, the missing asset may be a comparison, proof page, current FAQ, or product clarification. If the claim is false, content is not the fix.
Use this production contract:
Name the buyer question and split it into explicit conditions.
Attach each condition to a current claim, source, owner, date, and public URL.
Preserve the raw answer and note which conditions appeared, disappeared, or were misattributed.
Produce only what the diagnosed gap calls for.
Should every important claim live on one page?
No. The study supports condition tracking and evidence verification, not a universal consolidation rule.
One massive page can become stale, repetitive, or hard to own. Separate documentation is often the right home for security controls, integrations, regional availability, pricing, and legal terms. The production question is whether the evidence trail is explicit enough for a person or system to connect each condition to the same offer.
I checked the paper's limitations before turning its 5% evidence-delivery result into a content lesson. The two visual systems received credit only when surfaced pages matched stored reference pages. That is conservative. The paper also excludes proprietary retrievers and trained retrieve-and-verify agents.
Even with that caveat, the difference is hard to ignore. Conditional on retrieving a gold document, the two page-aware systems surfaced stored evidence for each individual condition about half the time. Full stored-page coverage fell to 13.6% and 17.2%. Across all 1,000 queries, Gold+Support reached only 5.1% and 5.3%.
Sometimes that means a comparison page that brings several facts together. Sometimes it means repairing a buried documentation link. Sometimes the evidence is missing and must be created. And sometimes the honest answer is that the company does not fit.
How can Trovance turn an incomplete answer into a production decision?
Trovance connects the observed answer to the weakest missing condition, then helps the team decide whether a proof-backed asset should exist.
A team can use Trovance to observe how AI systems explain, cite, compare, and recommend the company for a defined buyer question. The answer trail shows more than presence. It lets the team inspect whether each decision condition appeared, which source carried it, and whether a partial explanation made the company look like a fit for the wrong reason.
Trovance then helps diagnose the evidence gap behind the result. The missing condition may already have public proof that is stale, buried, or disconnected. It may require a comparison, product clarification, source-backed page, or no content action because the offer does not meet the constraint.
When production is justified, Trovance helps the team produce and publish the proof-backed asset that should exist, with claims, sources, and human review attached. A later rerun can test whether the explanation changed without promising a citation, recommendation, shortlist, or sale. Use Trovance to trace an incomplete answer back to the evidence decision.
FAQs
Does this study prove that AI search misses most website evidence?
No. CrossPage evaluates retrieval over long World Bank and PMC documents, not public websites or commercial answer engines. Its percentages describe one controlled instrument. The supported lesson is narrower: finding a relevant document and covering every requested condition are different events, so marketers should measure them separately in their own answer runs.
What is the difference between discovery and completion?
Discovery means a complete document appears somewhere in the top results. Completion means that document ranks ahead of known partial matches and can support every condition in the request. A system can discover the right source yet still lead with an incomplete option, which is why a hit rate alone can overstate performance.
What is a subset-first failure?
A subset-first failure occurs when a complete document is available in the top 10 but a document supporting only part of the request ranks ahead of it. In the study's strongest displayed hybrid, this happened for 45.3% of all queries, compared with 18.9% where no complete document appeared in the top 10.
Should teams combine all proof into one long page?
Not by default. Consolidation can help when several conditions belong to one buying decision, but separate documentation may remain the clearest source of truth. Audit the observed answer, source trail, ownership, freshness, and missing condition first. Then consolidate only when it improves truthful access without creating a stale or overloaded page.
What should a team do when proof exists but the answer omits it?
Check whether the proof is public, current, indexable, clearly connected to the company and offer, and written in language that resolves the buyer's condition. The right fix may be a stronger internal connection, clearer scope, or refreshed documentation. A new article is justified only when the missing asset has a distinct job.
How should multi-condition buyer questions be measured?
Record the model, date, full prompt or conversation, repeated-run policy, and each decision condition. For every run, mark whether the condition was stated, sourced, compared, and used in the recommendation. Keep the raw counts visible. A single score should never erase which condition failed or whether the company was a legitimate bad fit.
Can better content guarantee complete retrieval or a recommendation?
No. Content can make current proof easier to inspect and connect, but retrieval, synthesis, model behavior, competing evidence, and product fit remain outside a marketer's control. Test matched questions across repeated runs and report what changed. Do not turn a content update into a promise of citations, recommendations, pipeline, or revenue.
Related Resources
Trace the buyer question from observed answer to evidence action with Trovance

