ChatGPT gives first-person product picks in 79% of tests, arXiv audit says
An arXiv study audits AI shopping advice from ChatGPT and Google. It finds inconsistent recommendations, posing a measurement challenge for Q4 ad spend.
Hannah Vogel ·
In an arXiv preprint posted in September 2026, researchers audited how major AI systems answer real-world shopping questions and found that ChatGPT delivered first‑person product picks in 79% of product-recommending responses, versus 7% for Google Gemini and 2% for Google Search’s AI Overviews, across a dataset of 2,528 commercial-advice queries and 1,536 analyzed responses. The same study reports that recommendations and the sources cited often change across repeated requests, and that what APIs expose can differ materially from what consumer interfaces show. This is a single-source research preprint, not peer‑reviewed; the figures and claims are the authors’ and have not been independently verified by Atlas Media.
The assistants don’t agree on what to buy, and their sources barely overlap
The preprint’s central commercial finding is fragmentation. For the same query, the ChatGPT and Gemini consumer interfaces shared only 5.4% of domains on average, with no domain in common in 76.7% of comparisons. The APIs also diverged from their own consumer interfaces: mean domain overlaps were 12.0% for ChatGPT and 14.8% for Gemini. The authors add that product picks often changed across repeated requests, suggesting stochastic variation rather than a stable “top recommendation.” For operators who assumed an “assistant shelf” would act like the old ten-blue-links era of search, the audit points in the opposite direction: there is no single shelf to win and no single set of sources to optimize against. That undermines the premise of measuring or guaranteeing “share of recommendation” off a one‑time scrape or an API sample. [S1]
First-person recommendations are a branding asset for vendors — and a liability line for advertisers
The preprint highlights a stylistic difference with business implications: ChatGPT frequently uses a first‑person voice (“If I had to buy just one…”), while Google’s surfaces rarely do. If, as the authors note, OpenAI and Google are monetising their AI through advertising, the blend of a personal-sounding pick with paid placements will be scrutinized under endorsement and disclosure expectations. The paper does not test paid influence; it tests what the systems say today. But for marketers, the relevant point is practical: if your media plan includes assistant inventory, you should assume that the tone of the recommendation itself could shape user trust and conversion, and that each platform’s house style will change your creative and compliance work. A first‑person pick can read like an influencer spot; a generic overview reads like an aggregator. Those are different copy decks, different disclaimers, and different risk profiles for your legal team. [S1]
APIs are not a proxy for consumer reality — don’t buy tools that pretend they are
One under‑noticed result is how different the APIs look from the consumer products. The study reports low domain overlap between each vendor’s API and its own interface, and differences in “types and layers of source information” exposed. Developers love APIs because they are testable and scorable; procurement loves them because they scale. But if you are buying “assistant SEO” tools built only on API responses, the preprint’s message is simple: those tools are not measuring what your customers actually see. That has two commercial consequences. First, attribution models that attempt to link assistant recommendation exposure to sales need UI‑grounded panels, not just API logs. Second, performance guarantees tied to API rankings are mis-specified — they’re promising lift in a synthetic environment. Ask vendors to demonstrate repeat testing in the consumer interface, across multiple days, for your specific query set, and require they show variance bands, not just point estimates. The authors’ finding that “neither isolated responses nor API observations can be assumed to represent the commercial advice consumers encounter” is a direct warning against over-reliance on API‑only audits. [S1]
The dominant read — “optimize once for the assistant shelf” — misses the operating problem
A common reaction to assistant surfaces is to treat them like a new search results page: identify the featured sources, optimize content, negotiate affiliate links, and monitor share over time. The preprint’s variability data undermines that approach. If the same prompt yields different products and largely different sources on repeat — and if the API view doesn’t match the UI — then conventional SEO/affiliate playbooks will overfit noise. The operating problem here is measurement, not messaging. Brands planning Q4 and 2027 budgets should assume: 1) assistant exposure is volatile, 2) segment‑specific behavior matters (Gemini v. ChatGPT v. AI Overviews behave differently), and 3) the paid-organic boundary is evolving. That argues for pilot budgets with clear incrementality tests inside retailer or marketplace carts, not large “assistant optimization” retainers benchmarked to API‑based share of voice. [S1]
Channel power could shift to whoever controls the checkout, not the recommendation text
Because the preprint shows divergent recommendations and sources, the lever most likely to matter commercially is not what the assistant says but where it sends the user. If ChatGPT’s pick today links to a brand site and tomorrow to a publisher roundup, your traffic mix, affiliate costs and first‑party data capture will swing without your team changing anything. Google’s AI Overviews, by contrast, sit atop a search environment that already handles shopping journeys at web‑scale. The study does not measure click‑through or conversion, but its source‑overlap results imply that the distribution of traffic from assistants will be unstable. For brands, that means channel managers should emphasize end‑to‑end funnel instrumentation: UTMs, affiliate parameterization, and retailer attribution hooks — and be ready to adjust budgets away from assistant‑exposed content that cannot be tied to sales lift inside the cart. [S1]
The skeptic’s case: if recommendations vary, maybe their commercial impact is small
There is a reasonable counter‑read: if the systems don’t agree and answers shift on repeat, perhaps assistant advice is still peripheral to actual purchase behavior. The preprint does not present sales outcomes; it audits recommendations and sources. A marketer could look at a 5.4% domain overlap and argue that assistant optimization is premature until the platforms stabilize. That objection holds unless and until assistants begin to consistently insert shoppable links and native commerce flows. The audit’s point about differences between API and UI views still stands, however: if you are going to test at all, test what the customer sees, not what the developer sees. [S1]
What changes for marketers, agencies and procurement in the next two quarters
For CMOs and heads of media: treat assistant surfaces as an experimental channel with a measurement tax. Ring‑fence pilot budgets and define success as incremental sales in retailer or DTC carts, not shifts in “share of recommendation.” For agencies: avoid selling API‑only dashboards as truth; build repeatable UI‑capture rigs and disclose variance explicitly. For procurement: revise RFPs to require vendors to specify whether results come from API or consumer interfaces, to document the number of repeats per query, and to enumerate domains observed with timestamps. For legal and compliance: pre‑clear creative and disclosures for first‑person assistant environments separately from generic overview environments, because the preprint suggests users will encounter both. None of that depends on the platforms’ ad roadmaps; it depends on the study’s measured variability and stylistic differences, which are observable today. [S1]
Watch the labeling, the links and the testing standards
Three observable markers will tell operators whether this channel is maturing. First, policy updates: if OpenAI or Google publish ad or disclosure rules that explicitly cover chatbot product picks, the legal boundary is firming up. Second, link architecture: if assistants start consistently surfacing shoppable links or affiliate parameters in UI answers, brands can tie exposure to sales and justify budget reallocations. Third, research norms: if independent audits begin to report stabilized domain overlap or consistent “top picks” for common queries, or if vendors start referencing UI‑based, repeated testing standards in sales collateral, the market is moving from novelty to norm. The arXiv preprint’s core message — that neither isolated responses nor API observations reliably represent consumer reality — is the gating item. Until those markers move, treat this as an experiment you can instrument, not a shelf you can own. [S1]