AI shopping answers favor big retailers, not the stores they cite

Published:
October 7, 2026
Update:
October 7, 2026

AI shopping assistants are already acting like gatekeepers, and right now they tend to open the gate for big chains. In a 2026 Vaer AI study commissioned by Lightspeed Commerce, large national retailers dominated shopping recommendations in ChatGPT and Google’s AI experiences even when live web search was turned on. The most important lesson is not just bias toward scale. It is measurement: a smaller retailer can appear in the cited sources and still lose the actual recommendation the shopper sees.

  • Large retailers kept a clear edge in AI shopping answers, even with live search enabled.
  • Smaller merchants can be cited as sources without being recommended in the final answer.
  • Specific product queries widened the gap, while the word "independent" narrowed it far more than "local."
  • Retail GEO needs separate metrics for citations, mentions, and top recommendation share.

What did the study actually show?

The study found a consistent advantage for large retailers across both model-memory and live-search conditions. Vaer AI analyzed 20,000 shopping prompts across the U.S. and Canada, generating 200,000 answers without web search and 260,000 with search enabled. In the no-search condition, national chains accounted for about 63% to 70% of recommendations, while small or local independents appeared about 10% of the time. In a head-to-head choice between one large retailer and one smaller store, the models picked the larger retailer 90% to 94% of the time.

Live search helped, but only partially. With web search enabled, large retailers still captured 46% to 58% of top recommendations across ChatGPT, Google AI Mode, and Google AI Overviews. That matters because live retrieval is supposed to widen the field. Instead, it mostly softened the bias without removing it.

The gap widened as purchase intent got stronger. For broad shopping prompts such as "toys for a 6-year-old," small stores won roughly one-third of top recommendations. When the prompt named a specific product, their share fell to about 10%, while large retailers climbed from roughly 40% to 60%. In plain English, the closer the shopper gets to buying, the more the answer collapses toward the safest, most familiar merchants.

This does not automatically prove intentional favoritism by any one platform. What it does show is a repeatable outcome: when AI has to make a confident shopping choice, scale keeps beating specificity. For small merchants, that is not a branding problem alone. It is a distribution problem inside the answer layer itself.

Why does citation not equal recommendation?

Because AI answers have layers, and those layers do not behave the same way. A model can read a smaller retailer, cite it, and still recommend a national chain in the final sentence. That is exactly why teams that report only citations or only mentions can fool themselves about performance.

Viewed as a funnel, the study looked like this.

Layer in the AI answerWhat the study foundWhat it means for retailers
Source links citedLarge and small retailers were roughly even at about 38% eachYour brand can be present in the model’s evidence without winning the answer
Stores named in the answerLarge national chains still dominated the named recommendation layerMentions are more selective than citations
Lead recommendationLarge chains appeared about 2.5 times as often as small retailers as the top pickThe shortlist gets more concentrated as the model becomes more decisive

The distinction matters because it changes how you measure success. Among the stores cited as sources, large and small retailers were roughly even at about 38% each. But as the answer moved from source links to named stores to the single lead recommendation, big chains gained ground and small retailers lost it. By the lead pick, a large retailer won about 2.5 times as often as a small one.

That is the difference between visibility and preference. Visibility means the model found you relevant enough to read. Preference means the model trusted you enough to tell the user, "Buy here." In ecommerce GEO, those are separate stages with separate failure modes. A brand may be retrievable but not recommendable. It may be cited for background but not selected for action.

If you want a close parallel, think about AI visibility metrics that separate mentions, citations, and outcomes. The lesson is the same: one composite score can hide the exact point where you are losing. That is why raw mention counts are a weak KPI for shopping journeys with strong commercial intent.

BotRank's Take

This study lands on one of the biggest mistakes brands still make in Generative Engine Optimization (GEO): treating every appearance in an AI answer as if it meant the same thing. It does not. A brand that is cited in the supporting links, mentioned once in the body, and chosen as the top retailer has achieved three different outcomes.

That is why BotRank’s AI visibility tracking and Source Analysis matter in this context. Together, they help teams see which prompts produce a real recommendation, which sources models rely on, and whether the cited pages actually mention the brand they are supposedly supporting. For retail and marketplace brands, that distinction is practical, not academic. If you only track source presence, you can celebrate a visibility win while the model keeps sending shoppers to someone else. That is the kind of false positive that slows action and wastes content budget.

Why do specific shopping prompts help big retailers even more?

Because specificity increases the model’s need for confidence. When a user asks for a general category, the assistant can spread its answer across several merchant types. When the user asks for one exact product, the model seems to fall back toward merchants it already recognizes as broadly reliable and likely to stock the item. That explanation is an inference, but it matches the study’s pattern closely.

The prompt wording test makes this even more interesting. Adding the word "independent" more than doubled the share of small and local stores named, lifting them from about one-third of picks to nearly four-fifths in the randomized sample. By contrast, "local" and "near me" had much less effect, because a nearby chain can still qualify as local in the model’s logic. On Google’s platforms, the share of retailer sources coming from large national chains fell from about 44% under neutral prompts to as low as 9% when "independent" was added.

That single finding should change how retail marketers think about intent research. Users do not always know how to ask for the outcome they actually want. If your business depends on local discovery, the prompt space around "independent," "locally owned," or category-specific neighborhood intent may be more valuable than generic "near me" language. That is also why Prompts Studio is useful for structured testing across multiple models instead of one-off screenshots in a team Slack channel.

There is a second lesson here for content strategy. Broad category pages can still win discovery, but specific product and availability questions may require stronger proof signals, cleaner merchant data, and clearer page structure if you want the model to trust a smaller brand in a commercial moment. Put simply, recommendation bias becomes harder to fight as the question gets more precise.

What should retailers and ecommerce teams do now?

The answer is not "publish more content" and hope for the best. Retail teams should separate the recommendation problem into measurable layers, then work each layer on purpose.

  • Track recommendation share, not just mentions. If the business question is "Who does AI tell shoppers to buy from?" then top recommendation share deserves its own KPI. A flat mention count can look healthy while buyer-intent queries keep going to larger chains.
  • Build prompt panels by shopping stage. Keep broad discovery prompts separate from specific product prompts. The study showed that these stages behave differently, and your prompt monitoring should reflect that.
  • Test modifiers that reflect real buying intent. "Independent" performed differently from "local" and "near me." That is a reminder to validate language empirically instead of relying on intuition. BotRank’s Recommendations feature is helpful here because it turns analysis into concrete next actions instead of another slide deck.
  • Strengthen local and merchant trust signals across the open web. If you sell regionally, you need consistent identity, product clarity, and off-site validation, not just optimized category pages. Our guide on how local businesses win trust and visibility in AI search breaks down the signals that make smaller brands easier for models to trust.
  • Review model differences before you generalize. The same study showed that source behavior and recommendation behavior can vary across interfaces, even when the final outcome still favors big retailers. Treat ChatGPT, AI Mode, and Overview-style answers as related but separate environments.

One nuance matters. This playbook works well for merchants that can prove relevance, stock confidence, expertise, or local fit. It works less well when the business has thin product data, weak reviews, or inconsistent naming across its site and third-party sources. AI cannot recommend what it cannot understand, and it rarely trusts what it sees inconsistently.

There is also an opportunity hidden inside the study’s bad news. If large retailers are winning by default, then smaller brands do not necessarily need to be better known everywhere. They need to be better understood in the narrow contexts where intent is strong. That is an AI search visibility problem first, and a traffic problem second.

What does this change for ChatGPT and Google shopping strategy?

It changes the target from rankings alone to answer design. Shopping visibility in ChatGPT and Google’s AI Overview layer now depends on whether your brand survives three filters: being retrieved, being named, and being chosen. The study shows that many smaller merchants clear the first filter and fail at the third.

That means classic SEO still matters, but it is no longer the whole job. A page can be technically accessible, topically relevant, and even cited by the model, yet still lose because the answer engine has more confidence in a bigger merchant brand. For retail teams, the strategic question becomes simple: where does your brand drop out?

If it is never cited, you likely have a retrieval or discoverability problem. If it is cited but not named, you likely have a trust or entity problem. If it is named but rarely chosen first, you have a recommendation problem. Different layers, different fixes. That is the mindset shift this study should force.

FAQ

Does this prove that AI platforms intentionally suppress small retailers?

No. The study shows a strong and repeatable outcome, not proven intent. The safer reading is that model familiarity, retrieval patterns, and confidence in well-known merchants combine to favor larger brands.

Can a small retailer benefit from being cited even if it is not recommended first?

Yes, but it is only a partial win. The study shows that smaller retailers can appear in source links while still losing the main recommendation, which means citation alone should not be treated as commercial success.

Is "local" still a useful shopping modifier in AI prompts?

Sometimes, but it was much weaker than "independent" in this study. The reason is simple: a nearby branch of a national chain can still satisfy "local" in the model’s logic, while "independent" is harder to blur.

What should a GEO team measure first after reading this study?

Start with a fixed prompt set and split the outcomes into citations, mentions, and top recommendations across models. If you cannot tell which layer your brand is winning or losing, you cannot prioritize the right fix.

The bottom line is blunt. In AI shopping, being present is not the same as being picked. If you want to know whether your retail brand is actually winning inside commercial AI answers, run the prompts that matter, inspect the sources behind them, and measure who gets recommended first. BotRank gives teams a practical way to do that before "visibility" becomes another comforting metric that hides a revenue problem.

AI Search & GEO expert

After nearly 15 years in digital strategy on the client side (including 10 years at Olympique Lyonnais, where he was notably in charge of SEO).
Florian co-founded BotRank.ai in 2025, the GEO (Generative Engine Optimization) tool used by more than 2,500 companies to manage their visibility in AI-generated search results. He writes regularly about GEO and AI Search.