
11 min
Zach Chmael
In This Article
A 3,600-request simulation found cross-platform agents expanded product access while scarce shortlist attention shaped what buyers saw and chose.
Updated
TL;DR
๐งช The preprint's main comparison used 1,200 requests in each of three product domains.
๐ Its controlled access check moved target availability from 20.0% to 86.8% when an agent could query all six platforms.
๐ Expanding the market from 6 to 24 platforms raised access but lowered shortlist inclusion by 13.6 percentage points.
๐ At 24 platforms, the target was present in 97.7% of episodes but purchased in 37.9%.
๐งพ The study used 12 simulated users and one main random seed, so its rates are experimental results, not buyer benchmarks.
How Cross-Platform AI Agents Build Product Shortlists
A cross-platform AI agent can widen the pool of products it considers and still make attention harder to win. That is the useful result from a new 3,600-request simulation built from three Amazon review domains.
The agent carried one request across several simulated platforms, collected one candidate from each, and ranked a limited shortlist. Access improved because a relevant item could enter from any participating catalog. Yet as the market grew from 6 to 24 platforms, the target appeared more often while shortlist inclusion and purchase fell.
This is not evidence about live shopping agents or B2B revenue. It is a clean way to see the mechanism. Availability, candidate access, shortlist attention, first position, and purchase are separate states. "AI visibility" becomes useless when one score collapses them.
How does a cross-platform AI agent build a shortlist?
It starts with the buyer's request, queries participating platforms, collects their proposed products and explanations, and ranks a smaller cross-platform list.
That sequence differs from the recommendation system most people know. In a conventional marketplace, the buyer chooses a platform first. That platform controls the catalog and ranks items inside it. A relevant product outside the chosen catalog cannot compete.
The paper by Deyao Hong and five coauthors models a different path. A user agent interprets the request before platform choice, sends the same structured need to every platform, and receives one proposal from each. The agent then fills a shortlist with capacity K. The buyer can choose any shortlisted item or decline to purchase.

Source: Hong et al., The User Asks, Platforms Compete.
The diagram makes one thing plain: the agent becomes an allocator of attention. It does not merely search a larger catalog. Its query coverage determines who can enter.
Its shortlist determines who gets compared. Its ordering determines which candidate receives the first position.
The final buyer still decides. That matters because a first-place recommendation is not a purchase, and a purchase is not proof that the item was best.
What did the recommendation-market experiment test?
The experiment tested a synthetic market protocol over reconstructed retail requests, not real people delegating purchases to deployed agents.
The authors used Musical Instruments, Video Games, and Sports & Outdoors from Amazon Reviews'23. The underlying dataset work was published at ACL 2026 and evaluates semantic product retrieval over large review corpora. For this experiment, the authors retained users with at least 5 interactions, kept at most their 20 most recent, and held out a final item rated at least 4.
Only the query generator saw that held-out item. It wrote a plausible request for which the item would be a strong but non-unique fit. The platforms, ranking agent, and final decision model never received the target label.
This setup gives the experiment something concrete to trace. It also means every request is built around a known favorable reference. The resulting "Target Purchase" rate measures whether the simulated process recovered that item. It does not measure revenue, satisfaction, or the quality of every alternative.
The main comparison generated 1,200 requests per domain, balanced across six prompted request styles. The default market had six platforms, a three-item shortlist, catalogs capped at 30,000 regular items, and a target-inclusion probability of 0.8 when a catalog carried the target.
DeepSeek-V4-Flash handled the main generative roles. BAAI BGE-base-en-v1.5 handled dense retrieval. The runs used Gumbel temperature 0.1 and seed 42.
I downloaded the 38-page version-two PDF and checked the appendix because the neat diagram hides a lot of machinery. The main stateful comparison used 12 simulated users, two for each request profile. Those simulated continuities are not reconstructed Amazon customers.
Where does product access turn into scarce attention?
Access turns into scarce attention when the candidate pool grows faster than the shortlist.
The paper tracks the held-out target through four states: present in the candidate pool, included in the shortlist, ranked first, and purchased. The controlled protocol check puts the access difference in numbers. Under its default settings, forced target inclusion was 20.0% for platform-first recommendation and 86.8% for cross-platform mediation.
Natural retrieval could still recover the target in either condition. That is why those percentages are a sanity check for the market design, not an observed discovery rate.
In the main comparison, target presence rose from roughly one fifth of episodes under platform precommitment to nearly nine tenths under agent mediation. Target purchase was 4.2 to 4.8 times higher, mostly because the item gained a chance to compete at all.
Then the direction flipped. With shortlist capacity fixed at three, expanding the market from 6 to 24 platforms increased target presence by about 8.0 points. Shortlist inclusion fell about 13.6 points, and target purchase fell 13.0 points.

Source: Hong et al., The User Asks, Platforms Compete.
At 24 platforms, the target appeared in 97.7% of episodes. It reached the shortlist in 64.2% and was purchased in 37.9%. The gap between presence and shortlist inclusion was 33.5 percentage points.
The product did not become less available. The mean number of target copies more than doubled. Competing candidates grew faster, so the target's share of the pool nearly halved.
I checked the full figure before cropping it because "97.7% present" would make an irresistible and terrible benchmark. It is a domain mean from a controlled, target-conditioned simulation. The chart's real value is the divergence among stages.
Does a larger shortlist solve the attention bottleneck?
A larger shortlist creates more visible positions, but it does not create another first position.
The authors expanded shortlist capacity from 1 to 6 items. Target shortlist inclusion rose about 37.7 percentage points. Top-1 attention moved only about 1.1 points, and target purchase moved 2.2 points.
That result separates exposure from priority. A product can move from absent to visible while remaining below the position that frames the buyer's first impression. It can also reach first place and still lose the sale.
For a marketing team, this is the stage map worth preserving:
Stage | What it means | What it does not prove |
|---|---|---|
Available | The source or platform carries the product | The agent queried or retrieved it |
Candidate | The product entered the comparison pool | The buyer saw it |
Shortlisted | The agent exposed it in a limited set | It received first attention |
Top-1 | The agent ranked it first | The buyer accepted the recommendation |
Purchased | The simulated buyer selected it | The product caused satisfaction, retention, or revenue |
The paper adds a feedback mechanism after purchase, but even there the distinction matters. Its reputation memory tracks whether earlier recommendation claims matched a simulated later review. It does not reduce the outcome to a generic product rating.
In equal-supply mixed markets, promotional exaggeration held one third of candidate supply but captured 73% to 78% of first positions when the agent had no history. When a user-scoped agent could consult simulated outcome records, that share fell to 36% to 41%.
That part of the study is close to earlier work on persuasion in LLM recommendations, and it should stay bounded. The model wrote the explanations, ranking, decisions, reviews, and memory records. It shows what the protocol permits, not how real buyers behave.
What should marketers measure before claiming AI recommendation?
Marketers should measure each observable transition and stop naming all of them "visibility."
Start with one buyer question, one product condition, and a defined set of answer surfaces. Preserve the raw answer, cited sources, compared vendors, ordering, stated rationale, and any downstream action you can actually observe. Repeat the run because one answer is not a distribution.
Then ask where the brand left the path:
Diagnostic question | Possible evidence gap |
|---|---|
Was the brand eligible to be considered? | Product facts, entity identity, or access may be missing |
Did the system retrieve or cite the brand? | The relevant proof may be buried, stale, or attached to the wrong page |
Did the brand enter the comparison set? | The page may not state the buyer condition or differentiator clearly |
Was it recommended for the right reason? | The supporting claim may lack scope, proof, or a candid limitation |
Did a buyer act? | The gap may sit in product fit, price, trust, sales, or measurement rather than content |
Do not treat every loss as a publishing request. A missing source can justify a proof page. A stale comparison can justify a refresh.
A product limitation can require product work or a clearer boundary. Sometimes the right diagnosis is legitimate bad fit.
The preprint itself gives a good warning. Its purchase intervals use 50,000 bootstrap draws, yet the primary estimates still condition on one main model, one seed, one catalog assignment, and one episode order. More decimal places do not remove the experiment's boundary.

How can Trovance diagnose why a brand missed an AI shortlist?
Trovance can connect an observed answer or recommendation to the evidence transition where the brand disappeared.
A team can use Trovance to observe how AI systems explain, cite, compare, and recommend the company for a defined buyer question. The preserved answer trail shows whether the brand was absent, cited without being compared, compared for the wrong reason, or recommended without the claim a buyer would need to verify.
A visibility score alone cannot distinguish those failures. The underlying gap might be a missing product fact, weak comparison evidence, an inaccessible source, stale proof, or legitimate bad fit. Trovance helps diagnose that evidence gap before the team commits to another asset.
When publication is justified, Trovance helps produce and publish the proof-backed page, comparison, FAQ, or source-backed article that should exist. A person keeps responsibility for truth, fit, and publication. A later rerun can check whether the explanation changed without promising a citation, shortlist, sale, or revenue result. Use Trovance to trace an AI shortlist gap back to the evidence your team can repair.
FAQs
What is an agentic recommendation market?
An agentic recommendation market is a proposed setup where a user asks an AI agent before choosing a platform. The agent queries participating platforms, collects candidate products and explanations, and ranks a cross-platform shortlist. The paper studies this as a controlled protocol, not as proof that current shopping agents work this way in production.
Does product availability mean an AI agent will recommend it?
No. Availability only means a platform or source carries the product. The agent still has to query that source, retrieve the item, admit it to the candidate pool, place it on a limited shortlist, and rank it high enough to matter. Each transition can fail for a different reason.
Did this study use real buyers or live shopping agents?
No. It used LLMs to generate requests, platform explanations, rankings, purchase decisions, reviews, and reputation updates. The requests were grounded in Amazon interaction histories, but the sessions and user continuities were simulated. The results explain a mechanism inside the testbed, not measured behavior from ChatGPT, Gemini, or real consumers.
Why did adding more platforms reduce target purchase?
Adding platforms increased the chance that the held-out target entered the candidate pool, but it increased competing proposals even faster. With the shortlist fixed at three items, the target's share of the pool fell. More access therefore created more competition for the same small number of visible positions in this experiment.
Would a longer AI shortlist give brands more attention?
It would give more candidates some exposure, but the study found that most of the gain appeared below first place. Expanding capacity from one to six items raised target shortlist inclusion sharply while Top-1 attention and target purchase changed little. A longer list creates slots, not another first position or a guaranteed decision.
Can reputation history stop exaggerated recommendations?
The simulation found that user-scoped outcome history reduced exaggerated explanations' share of first positions, but it did not eliminate them. The history compared past claims with simulated post-purchase reviews. A deployed system would face sparse evidence, cold starts, changing products, strategic behavior, privacy questions, and uncertainty that this controlled testbed does not resolve.
What should a B2B team do after finding a shortlist gap?
Preserve the exact buyer question, answer, sources, compared vendors, ordering, and rationale. Diagnose whether the loss came from identity, relevance, proof, comparison language, access, product fit, or reputation. Publish only when a public evidence asset can address the gap. Otherwise refresh, clarify, seek third-party proof, fix the product, or take no action.
Related Resources
Trace the missing shortlist stage back to a defensible evidence action with Trovance

