AI Chatbots Dispense Risky Health Advice
AI chatbots in a BMJ Open study gave problematic medical answers nearly half the time, raising public health concerns as use grows.
Ayla Demirhan ·

Artificial intelligence chatbots including Google's Gemini, DeepSeek, Meta AI, ChatGPT, and Grok produced inaccurate or incomplete answers to medical questions in a study published Tuesday in BMJ Open . Researchers said nearly half of the responses they reviewed were classified as “problematic,” raising concerns about how people may use these tools when seeking health information.
The research team at the Lundquist Institute for Biomedical Innovation at Harbor-UCLA Medical Center tested the systems with prompts designed to “strain” them and draw out misinformation. The study focused on topics such as cancer, vaccines, and nutrition, areas where misleading guidance can carry significant consequences for patients and public health.
According to the study, 30% of chatbot responses were rated “somewhat problematic,” while 19.6% were assessed as “highly problematic.” Researchers said the issues included answers that were wrong, incomplete, or framed in ways that could confuse users about what is supported by evidence.
The authors reported that even when a chatbot warned against unscientific treatments, it often still presented alternative therapies such as acupuncture or herbal medicine. In some cases, the systems identified clinics offering such approaches, including Gerson therapy, which discourages chemotherapy. The researchers described this pattern as “false balance,” meaning scientific and unscientific claims were presented with similar weight.
Performance varied by topic and by model. The study said the chatbots were most accurate on vaccine and cancer questions, but it also found that more than a quarter of cancer-related responses were potentially harmful. Among the tested models, Grok showed the lowest performance, the researchers reported.
The findings were framed as a public health risk in part because of how widely these tools are used. The study noted that approximately one-third of adults use AI for health information, which can amplify the impact of misleading or incomplete answers when users treat chatbot output as guidance rather than general information.
For policymakers, healthcare providers, and technology companies, the study underscores a continuing challenge: AI systems can generate confident-sounding responses that mix evidence-based medicine with unsupported options. Researchers said this can contribute to patients choosing unproven treatments over conventional care, particularly in high-stakes areas such as cancer.
While the study documents the scale of problematic answers under testing conditions designed to elicit misinformation, it does not resolve how quickly model behavior may change as systems are updated. The authors’ results nonetheless highlight uncertainty for users and regulators about the reliability of AI-generated health information across platforms and over time.