Errors, Hallucinations, and Clinical Impact of General-Purpose Multimodal Large Language Models in Histopathology
Thursday, September 24, 2026
Background: General-purpose large language models (LLMs) are increasingly evaluated in diagnostic pathology, but prior studies have largely emphasized diagnostic accuracy rather than how models fail. We evaluated four LLMs for diagnostic performance, pathology-relevant errors and hallucinations, their burden, and potential clinical impact across multi-organ pathology cases. Design: In this retrospective multicenter study, 153 pathology cases from two institutions spanning 20 organs were evaluated using ChatGPT-5.3 (LLM1), Gemini 3 (LLM2), Grok 4.20 (LLM3), and Claude Opus 4.6 (LLM4). Each LLM received multi-magnification histologic images with clinical context and generated a microscopic description and diagnosis. No data splitting, training, or fine-tuning was performed. Twenty-one pathologists assessed 612 outputs for diagnostic correctness, error and hallucination type and burden, clinical impact (0-4), and overall performance (1-5). Results: Strict diagnostic accuracy was 48.9% overall (60.1% including partially correct diagnoses) and ranged from 36.6% to 58.2% across LLMs. Errors occurred in 82.0% of outputs and hallucinations in 76.6%; 90.2% contained at least one error or hallucination. Misinterpretation was the most frequent error (72.1%), while fabricated histologic features were the dominant hallucination type (74.5%). LLM2 had significantly lower misinterpretation rates than the other three LLMs, while LLM3 had significantly higher fabricated-feature hallucination rates than LLM1 and LLM2. Errors were more frequent in incorrect than correct diagnoses (97.5% vs 66.2%), as were hallucinations (95.1% vs 59.9%; both p < 0.001). Every strictly incorrect diagnosis contained an error and/or hallucination, while 79.9% of strictly correct diagnoses also contained at least one. In multivariable analysis, LLM3 was independently associated with higher odds of an incorr...
Daily healthcare AI brief
Get the free healthcare AI briefing
One concise briefing built from reporting, research, policy, blogs, videos, and full podcast transcripts.
Free. No account needed. Unsubscribe anytime. Already subscribed? Sign in to save and personalize.