ArXiv study says AI assistants misfire on local referrals without search, changing who wins leads
A new study audits how AI assistants recommend local professionals. Without web search, results are often inaccurate and biased toward firms with misconduct.
Hannah Vogel ·

In a September preprint on arXiv, researchers audited how AI assistants recommend local providers across four registry-backed service categories, matching every recommendation against Medicare clinician and facility records and SEC adviser disclosures for the 100 largest U.S. metropolitan areas. The study tested three configurations: an open‑weight model, a proprietary model without web search, and the same proprietary model with search, then measured whether each recommendation matched a real provider in the queried city. The authors report that without search, only 4% of the open‑weight model’s recommended doctors and 11% of the proprietary model’s matched a clinician in the right city; with search enabled, 64–71% of recommendations matched a real provider in the same domains. The paper also finds that search flips which firms are recommended and reduces a bias toward advisers with SEC misconduct disclosures. This is, so far, single‑source — an arXiv preprint — with no independent confirmation, and no one in the reported packet is on the record.
A referral channel that looks like lead-gen — and carries compliance risk when retrieval is off
The paper’s core claim is blunt: when AI assistants answer provider questions without retrieval, they often fabricate referrals in domains the web covers thinly. In healthcare, the open‑weight model’s rare “matches” were largely name coincidences and were no likelier to be primary‑care clinicians than a random draw from the Medicare registry. In financial advice, the proprietary model without search more often recommended firms with SEC misconduct disclosures, at 3.6 times the base rate in the registry even after adjusting for firm size. When search is turned on, the match rates jump to the mid‑60s to low‑70s and the misconduct skew flips — recommendations fall below the registry’s disclosure prevalence. If you operate in regulated referral flows — payers steering to doctors, banks suggesting advisers — the finding is not an AI novelty story. It’s a distribution and liability story: the assistant’s retrieval configuration determines whether the “lead” you hand a customer is a real, qualified provider or a hallucinated name that raises exposure.
Search doesn’t just validate; it reallocates demand across firms and cities
The study argues that enabling search does more than check facts. It changes who gets recommended, and where. Without search, “real” matches cluster in the largest metros — a metro‑size penalty for smaller markets. With search, those match rates become similar across metro‑size terciles. In practical terms, an assistant that relies on its pretraining is over‑indexed to big‑city references and brand familiarity; a search‑augmented assistant redistributes visibility toward providers whose information is actually retrievable in smaller markets. For marketing leaders in multi‑location services, that is a channel design question. If the AI surface your customers increasingly consult is retrieval‑driven, then the completeness and crawlability of your location data — hours, specialties, licenses — changes from hygiene to revenue allocation. If it is not retrieval‑driven, the paper suggests the winners are whoever’s namesakes were most widely mentioned online years ago, whether or not they practice in a customer’s city.
Restaurants expose the visibility premium in AI outputs more than the quality one
In the restaurant domain, where quality (ratings) and visibility (review volume) can be measured separately, the study finds a 3–5x review‑count premium in recommendations but at most a tenth‑of‑a‑star rating premium. That implies assistants weight scale and salience far more than marginal differences in consumer‑reported quality when making referrals, even when search is present. If that pattern generalizes to other local categories, “being the most reviewed” matters more than “being marginally higher rated” for inclusion in AI answers. For marketers, that recasts the ROI calculus: initiatives that drive verifiable, crawlable signals of presence (structured listings, consistent NAP data, rich service pages) may have more impact on AI referral inclusion than incremental reputation management aimed at nudging a 4.3 to a 4.4. The paper does not claim causality beyond the observed correlations, but the operational implication is clear enough: assistants privilege what they can confidently find and summarize.
Procurement will need to contract for retrieval defaults and registry grounding in regulated journeys
The paper highlights a governance gap: “an answer produced without retrieval often carries no sign that its recommendations were never verified.” If an insurer, health system, or wealth manager embeds an assistant into a customer journey and omits retrieval or registry grounding, the default UX may not signal that the names offered weren’t checked against official records. That becomes a procurement item, not a model‑selection debate. Buyers should be specifying, and testing, default on‑state retrieval for regulated verticals, explicit grounding to the relevant registries (Medicare for clinicians, SEC for advisers), and visible citations or labels that surface the provenance of a recommendation. The research’s 3.6x misconduct‑disclosure skew absent search is a concrete reminder that “hallucination” can map to compliance exposure, not just trivia errors. Absent clear defaults and audit logs, you are effectively outsourcing your referral steerage to a pretraining corpus with unknown recency and coverage.
The obvious counter: defaults are changing fast; pilots aren’t production — but the audit’s denominator test still bites
Assistant vendors will argue that modern deployments default retrieval on for most real‑world use cases, and that enterprise configurations increasingly include RAG over vetted corpora rather than open‑web search. They may also note that an arXiv preprint is not peer‑reviewed and that prompt design, safety filters, and model versions can materially change outcomes. Those points can all be true and still leave the denominator test standing: in the audit’s setup, without retrieval, only a low‑double‑digit share of recommendations matched a real provider in the right city in thin‑web domains, while enabling search lifted matches into the 60s and 70s and altered who was recommended. If defaults and enterprise patterns have improved since the study’s runs, that should be observable in assistant UIs (labeling, citations) and in independent replications, not assumed.
For local-services marketers, the spend question is where budget physically moves next year
If assistants become a material referral surface, budget will follow. The paper’s findings point to two practical shifts. First, dollars move from brand lift and review‑score polishing to structured visibility: schema markup, verified listings synced to official registries where they exist, and content that aligns with how assistants extract attributes (specialties, accepted plans, languages). Second, in regulated categories, go‑to‑market teams will need closer ties to risk and compliance to ensure that any AI‑mediated steerage in owned channels is retrieval‑on and registry‑grounded by default — and that partner platforms that generate leads attest to the same. The restaurant result suggests that “how many verifiable signals do we have?” may matter more than “are we 0.1 stars higher?” in whether an assistant includes you. Expect RFPs for local‑presence management to cite assistant coverage explicitly, and for providers in smaller metros to compete more evenly if retrieval is table stakes.
What to watch in the next two quarters: defaults, labels, and lead-gen mix
Three signals will show whether this research describes a passing artifact or a durable channel rule. First, watch whether mainstream assistants begin labeling provider answers with explicit citations to registries in healthcare and finance by default — a visible proxy for retrieval‑on plus authoritative grounding. Second, in performance data from local‑presence vendors and marketplaces, look for a shift in the source mix of inbound leads that teams can attribute to AI assistant exposure; if search‑backed assistants rise, multi‑location providers should see more even geographic distribution of AI‑sourced leads across metro sizes. Third, in enterprise procurement, expect to see contract language that specifies retrieval defaults and grounding sources for any assistant deployed in regulated referral flows; if those clauses appear, the compliance reading of this paper has landed in buying behavior. If none of these appear by mid‑2027, the study may have captured an early configuration gap that vendors have since closed.
The arXiv preprint makes a simple, testable claim: retrieval configuration, not just the underlying model, drives whether an AI referral is a trustworthy steer in local services. For operators, that reframes AI from a model‑selection problem to a channel and contracting problem. The marketing work moves to being findable in the structures assistants use; the buying work moves to enforcing retrieval and registry grounding where it matters; and the governance work moves to making sure customers can see when answers are actually verified.