Image-Based Diagnosis of Oral Lesions: Performance of a Vision-Language Model versus Human Clinicians.
Monday, September 21, 2026
Published in Otolaryngology--head and neck surgery : official journal of American Academy of Otolaryngology-Head and Neck Surgery. Abstract: To evaluate the real-world diagnostic performance of a multimodal large language model (LLM) for image-based assessment of oral mucosal lesions compared with clinicians of varying expertise. Prospective international multicenter diagnostic accuracy study. Twenty university and tertiary head and neck centers in Italy, Belgium, France, Spain, and Israel. We enrolled 350 consecutive patients (320 with oral lesions, 30 with normal mucosa). Clinical photographs and basic epidemiologic data were analyzed using Gemini 2.5 Advanced with a standardized prompt. Model outputs for lesion detection, malignancy versus benign versus normal, precise histologic diagnosis, and urgency class were compared with histopathology and with 4 clinicians. Sensitivity, specificity, accuracy, and agreement were calculated. AI-Gemini achieved 97.1% accuracy for lesion detection and malignancy classification, with sensitivity 98.5% and specificity 96.2% for malignancy, and 88.0% accuracy for precise histologic diagnosis. The head and neck surgeon achieved the highest accuracy for precise diagnosis (97.7%). Three-class diagnostic accuracy was 94.2% for AI-Gemini and 67.0% to 86.5% for nonsurgeon clinicians. Urgency assignment was correct in 70% of cases (κ = 0.716), with a conservative tendency to overestimate risk. Agreement with histologic diagnosis was almost perfect (κ = 0.929). In this exploratory study, a multimodal LLM showed encouraging performance in image-based evaluation of oral mucosal lesions. ...
Daily healthcare AI brief
Get the free healthcare AI briefing
One concise briefing built from reporting, research, policy, blogs, videos, and full podcast transcripts.
Free. No account needed. Unsubscribe anytime. Already subscribed? Sign in to save and personalize.