A firm deciding how much attention to give answer engines will find the question apparently settled, twice, in opposite directions.
The disagreement
An analysis of 973 ecommerce sites covering roughly 20 billion dollars of revenue between August 2024 and July 2025, reported by Search Engine Land, found that traffic referred by large language models converted worse than Google search, email and affiliate links, on both conversion rate and revenue per session.
Other published measurements point the other way. ThoughtMetric reported 6.7 percent conversion for language model traffic against 3.9 percent for organic search. Another reported 15.9 percent against 1.8 percent. Semrush has estimated value per session from language model referrals at roughly 4.4 times organic search. A peer-reviewed study of ChatGPT referrals to ecommerce sites appeared in Marketing Science.
These are not all measuring the same thing. Sample composition, the mix of categories, whether branded search is included in the organic baseline, and how a session is attributed all move a conversion comparison by more than the gap being argued about.
What is not in dispute
Volume is small and growing fast. In the ecommerce dataset, referrals from language models were about 0.2 percent of sessions, roughly two hundred times smaller than Google organic. Over 90 percent of that came from ChatGPT alone. Semrush clickstream analysis put year on year growth in outbound referrals at around 206 percent comparing January 2025 with January 2026.
A channel at 0.2 percent of sessions growing at that rate is not yet a revenue line and is not safe to ignore either. Both of those are true at the same time, and most of the advice being published resolves the tension by picking one.
What we take from it
The practical answer is to stop treating the conversion question as something to resolve by reading. Your own referral data from these sources is measurable now, at no cost, and it answers the only version of the question that affects your decisions.
The wider point is about how to read a market in motion. When credible studies disagree by a factor of five, the disagreement is the finding. It means the effect is real enough to measure and unstable enough that anyone quoting a single figure with confidence has either not read the others or has a reason not to mention them.
We hold this as a working thesis rather than a position: presence inside answer engines is worth building for the placement rather than for the click, and the traffic case for it is not yet settled either way.