MEBRO

DISINFO DESK

Science & Health

AI Models Spread Medical Misinformation: Mount Sinai Study Reveals 32-47% Error Rates

Peer-reviewed Lancet study reveals leading medical AI models accept false health claims 32-47% of the time. 40M daily users, documented harm cases in patient-safety testing, but simple safety prompts cut errors nearly in half. Full analysis of Mount Sinai's 3.4M prompt test across 20 LLMs.

TRUE

FILED AUG 15, 2026 · UPDATED AUG 15, 2026 · 20 SOURCES

The Study: 3.4 Million Prompts, 20 AI Models, Alarming Results

The findings exposed a critical vulnerability called "sycophancy"—AI models' tendency to agree with authoritative-sounding content regardless of factual accuracy.[17] When false claims appeared in hospital discharge notes, AI acceptance rates jumped to 47%, versus only 9% for social media posts.[6] Models were particularly susceptible to two rhetorical tricks: appeals to authority ("a senior doctor says this") at 34.6% acceptance, and slippery slope arguments at 33.9%.[7]

Performance varied dramatically by model: ChatGPT-4o showed the strongest resistance at 10.6% susceptibility, while smaller models like Gemma-3-4B-it accepted false claims 63.6% of the time.[5][18] Shockingly, medical fine-tuned models consistently underperformed general-purpose models across all tests.[8]

False Claims AI Models Believed

The study documented AI models accepting demonstrably false claims including:[5][19]

In one alarming example, a discharge note falsely advised patients with esophagitis-related bleeding to "drink cold milk to soothe the symptoms." Several models accepted the statement rather than flagging it as unsafe.[4]

Source Credibility Paradox: AI Trusts Doctors More Than Reddit

Counter-intuitively, AI models demonstrated dramatically different susceptibility based on how information was presented. When misinformation came from what looked like an actual hospital note from a healthcare provider, the chances that AI tools would believe it and pass it along rose from 32% to almost 47%. Interestingly, AI was more suspicious of social media—when misinformation came from a Reddit post, propagation by AI tools dropped to just 9%.[6]

Co-senior author Dr. Eyal Klang explained: "Our findings show that current AI systems can treat confident medical language as true by default, even when it's clearly wrong. A fabricated recommendation in a discharge note can slip through. It can be repeated as if it were standard care. For these models, what matters is less whether a claim is correct than how it is written."[2][3]

The Scale of Exposure: 40 Million Daily Users

The urgency of these findings is amplified by the massive scale of AI health information seeking. According to OpenAI's January 2026 report, more than 40 million people ask ChatGPT healthcare questions every day, signaling consumers are frequently turning to the chatbot to navigate the complex healthcare system.[9] More than 5% of all ChatGPT messages globally are about healthcare, averaging billions of messages each week.[10]

Approximately 70% of health-related ChatGPT conversations occur outside standard clinical hours, indicating people turn to AI when they cannot reach healthcare providers. Rural communities and hospital deserts show particularly heavy reliance, with approximately 600,000 healthcare messages per week from rural areas and 580,000 weekly messages from areas more than 30 minutes from a hospital.[10]

Documented Patient Harms: From Theory to Reality

ECRI's 2026 Health Technology Hazards report named AI chatbot misuse the number one patient safety threat for 2026. The nonprofit patient safety organization documented multiple cases of harm:[11][12]

Dr. Marcus Schabacker, ECRI President and CEO, has explained the risk this way: it isn't that chatbots have suddenly become more dangerous, but that when an AI answer feels helpful and definitive, people start relying on it without questioning it.[11]

Hope on the Horizon: Simple Safety Prompts Cut Errors in Half

However, in the companion Communications Medicine study, Mount Sinai researchers discovered a simple intervention that dramatically improved safety. Adding a single-line "safety prompt" warning models that input information might be inaccurate reduced hallucination rates from 66% to 44%—and dropped ChatGPT-4o's error rate from 53% to just 23%.[13][14]

Lead author Dr. Mahmud Omar stated: "Hospitals and developers can use our dataset as a stress test for medical AI. Instead of assuming a model is safe, you can measure how often it passes on a lie, and whether that number falls in the next generation."[19]

The study represents a paradigm shift in AI safety evaluation—treating "can this system pass on a lie?" as a measurable property that must be stress-tested before clinical deployment. While the vulnerability is real and documented, the research also demonstrates that targeted safeguards can significantly reduce risks, making safe medical AI achievable with deliberate engineering and human oversight.

The Medical Fine-Tuning Paradox

One of the study's most surprising findings challenges a core assumption in AI development: models specifically trained on medical data performed worse than general-purpose models across all tests.[5] This counter-intuitive result suggests that current medical fine-tuning approaches may reduce general reasoning capabilities while adding medical vocabulary, without improving the fundamental ability to distinguish truth from confident-sounding falsehood.

The finding does not mean medical AI is impossible, but rather that current approaches need refinement. Post-training safety alignment appears to be more important than model size alone, with GPT-4o's strong performance (10.6% susceptibility) demonstrating that careful engineering can produce safer systems.

Rhetorical Tricks That Fool AI

The study systematically tested how rhetorical framing affected AI susceptibility. Models accepted 34.6% of fake claims that included the phrase "an expert says this is true," and 33.9% of fake statements when prompted with "if X happens, disaster follows."[7] These findings reveal specific vulnerabilities that can be targeted for improvement.

Discussing the broader pattern the research team observed across their studies, Dr. Mahmud Omar noted: "What we saw across the board is that AI chatbots can be easily misled by false medical details, whether those errors are intentional or accidental. They not only repeated the misinformation but often expanded on it, offering confident explanations for non-existent conditions."[13]

Methodology: Two Complementary Studies

The February 2026 Lancet Digital Health study built on a previous Mount Sinai study published in Communications Medicine in August 2025. Together, these studies provide comprehensive evidence of AI medical misinformation vulnerability:

Communications Medicine Study (August 2025): Researchers created 300 physician-designed clinical scenarios with fabricated medical details, including a fictitious diagnosis ("Casper-Lew Syndrome"), a fake lab test ("serum neurostatin"), and a made-up symptom ("cardiac spiral sign"). They tested 6 major LLMs to measure "hallucination" rates—how often AI elaborated on nonexistent medical conditions. Baseline hallucination rates ranged from 50-83% across models without safeguards.[15][16]

The Lancet Digital Health Study (February 2026): Expanded research to 20 LLMs across 3.4 million prompts using three realistic data sources: MIMIC hospital database discharge notes (with one fabricated recommendation inserted), Reddit health myths, and 300 physician-validated scenarios. The study systematically varied rhetorical framing to identify specific vulnerabilities.[7][8][18]

Solutions: A Roadmap for Safer Medical AI

The researchers and patient safety organizations recommend a multi-layered approach to medical AI safety:

Dr. Girish N. Nadkarni, Mount Sinai's Chief AI Officer, emphasized: "AI has the potential to be a real help for clinicians and patients, offering faster insights and support. But it needs built-in safeguards that check medical claims before they are presented as fact. Our study shows where these systems can still pass on false information, and points to ways we can strengthen them before they are embedded in care."[2][3]

Institutional Authority: Mount Sinai's Windreich Department

The research was conducted by Mount Sinai's Windreich Department of AI and Human Health, the first department of its kind at a U.S. medical school.[2] Leadership includes Dr. Girish N. Nadkarni (Chair and Chief AI Officer), Dr. Eyal Klang (Chief of Generative AI), and Dr. Mahmud Omar (physician-scientist and first author of both studies).[2]

The study was published open access in The Lancet Digital Health with full methodology transparent and reproducible.[1] The DOI is 10.1016/j.landig.2025.100949, and authors include researchers from Mount Sinai, Mayo Clinic Department of Radiology, and the Hasso Plattner Institute for Digital Health.[1]

GenuVerity Assessment

What is TRUE: Independent, peer-reviewed research (The Lancet Digital Health, Feb 2026) confirms leading AI models accept false medical claims at meaningfully high rates — 32% overall, rising to roughly 47% when misinformation is dressed up as a hospital discharge note.[1][6][7]

What Provides HOPE: A single-line safety prompt cut hallucination rates nearly in half in Mount Sinai's companion study, and the strongest model tested (ChatGPT-4o) resisted misinformation far better than smaller or medically fine-tuned models — evidence that safer medical AI is achievable with deliberate engineering.[13][14][5]

What is MISINTERPRETATION: The findings don't mean AI models are uniformly unreliable — susceptibility varies enormously by model and by how information is framed. The study measures a controlled propensity to repeat planted misinformation in testing, not a tally of real-world patient injuries.[5][11]

Harm Potential: HIGH (but mitigatable) — 40 million daily users, documented harm cases in patient-safety testing, but technology is improvable.[9][10][11] This is a legitimate and urgent patient safety concern that requires immediate action on multiple fronts: regulatory frameworks, institutional governance, user education, and continued AI safety research. However, it is not a reason to abandon medical AI—rather, a call to deploy it responsibly with appropriate safeguards and human oversight.

SOURCES · 20

  1. [1]The Lancet Digital Health - Mount Sinai Study — thelancet.com

    95/100 · thelancet.com

  2. [2]Mount Sinai Newsroom - Official Press Release — mountsinai.org

    72/100 · mountsinai.org

  3. [3]EurekAlert - Academic Press Release — eurekalert.org

    72/100 · eurekalert.org

  4. [4]Medical Xpress - Research Coverage — medicalxpress.com

    72/100 · medicalxpress.com

  5. [5]Euronews Health - International Coverage — euronews.com

    82/100 · euronews.com

  6. [6]HIT Consultant - Health IT Analysis — hitconsultant.net

    72/100 · hitconsultant.net

  7. [7]Digital Information World - Tech Coverage — digitalinformationworld.com

    72/100 · digitalinformationworld.com

  8. [8]Inside Precision Medicine - Clinical Analysis — insideprecisionmedicine.com

    72/100 · insideprecisionmedicine.com

  9. [9]Fierce Healthcare - OpenAI Usage Report — fiercehealthcare.com

    72/100 · fiercehealthcare.com

  10. [10]Healthcare Dive - ChatGPT Health Usage — healthcaredive.com

    72/100 · healthcaredive.com

  11. [11]ECRI 2026 Health Tech Hazards Report — ecri.org

    72/100 · home.ecri.org

  12. [12]MedTech Dive - ECRI Patient Safety Report — medtechdive.com

    72/100 · medtechdive.com

  13. [13]AzoAI - Safety Prompt Study — azoai.com

    72/100 · azoai.com

  14. [14]Mount Sinai - Communications Medicine Study — mountsinai.org

    72/100 · mountsinai.org

  15. [15]Medical Economics - AI Skepticism Study — medicaleconomics.com

    72/100 · medicaleconomics.com

  16. [16]U.S. News - Fake Medical Terms Test — usnews.com

    72/100 · usnews.com

  17. [17]Creati.ai - Sycophancy Analysis — creati.ai

    72/100 · creati.ai

  18. [18]Heise Online - Model Performance Analysis — heise.de

    72/100 · heise.de

  19. [19]News-Medical.Net - Study Coverage with Researcher Quotes — news-medical.net

    72/100 · news-medical.net

  20. [20]Bioengineer.org - Safeguards Analysis — bioengineer.org

    72/100 · bioengineer.org

MEBRO · DISINFO DESK · mebro.app

Investigative report — not a user-submitted fact-check.

AI-built, source-verified. Every claim here was checked against the sources cited above before publishing — but don't just trust us: follow any citation to its source and confirm it yourself. That's the whole point.