Introduction Psoriasis is common, but diagnosis and early severity assessment can be delayed because of variable presentation and overlap with mimicking dermatoses. Multimodal large language models (LLMs) may assist image-based triage.Objective To evaluate web-based multimodal LLMs for psoriasis identification, Physician Global Assessment (PGA) scoring, and treatment recommendation quality from clinical photographs.Methods We retrospectively analyzed 303 standardized photographs from 160 patients (Semmelweis University, May 2022-January 2025), including 163 psoriasis lesions and 140 mimickers. Reference diagnosis and PGA were assigned by two dermatologists with third-expert adjudication, treatment outputs were rated for appropriateness. ChatGPT-5, ChatGPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.5 received identical prompts for diagnosis and, when psoriasis was predicted, PGA and treatment; sessions were periodically reset.Results Diagnostic accuracy was highest for ChatGPT-5 (93.1%) and ChatGPT-4o (90.1%), followed by Claude (83.6%) and Gemini (61.5%). Among correctly identified psoriasis cases, PGA accuracy was 93.3% (ChatGPT-5), 92.9% (ChatGPT-4o), 82.7% (Gemini), and 78.1% (Claude). Appropriate treatment recommendations were most frequent for ChatGPT-5 (83.7%) and ChatGPT-4o (82.1%), then Claude (75.0%) and Gemini (53.2%); test-retest agreement favored the OpenAI models.Conclusion Under standardized conditions, general-purpose multimodal LLMs, especially ChatGPT-5 and ChatGPT-4o, showed strong performance for psoriasis recognition and reasonable support for PGA scoring and treatment suggestions, supporting potential use by primary care physicians when dermatology specialist access is limited.
Evaluation of multimodal large language models for psoriasis diagnosis, severity grading, and treatment recommendations from clinical photographs: ChatGPT shows superior performance compared to other large language models
Gisondi, Paolo;Bellinato, Francesco;
2026-01-01
Abstract
Introduction Psoriasis is common, but diagnosis and early severity assessment can be delayed because of variable presentation and overlap with mimicking dermatoses. Multimodal large language models (LLMs) may assist image-based triage.Objective To evaluate web-based multimodal LLMs for psoriasis identification, Physician Global Assessment (PGA) scoring, and treatment recommendation quality from clinical photographs.Methods We retrospectively analyzed 303 standardized photographs from 160 patients (Semmelweis University, May 2022-January 2025), including 163 psoriasis lesions and 140 mimickers. Reference diagnosis and PGA were assigned by two dermatologists with third-expert adjudication, treatment outputs were rated for appropriateness. ChatGPT-5, ChatGPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4.5 received identical prompts for diagnosis and, when psoriasis was predicted, PGA and treatment; sessions were periodically reset.Results Diagnostic accuracy was highest for ChatGPT-5 (93.1%) and ChatGPT-4o (90.1%), followed by Claude (83.6%) and Gemini (61.5%). Among correctly identified psoriasis cases, PGA accuracy was 93.3% (ChatGPT-5), 92.9% (ChatGPT-4o), 82.7% (Gemini), and 78.1% (Claude). Appropriate treatment recommendations were most frequent for ChatGPT-5 (83.7%) and ChatGPT-4o (82.1%), then Claude (75.0%) and Gemini (53.2%); test-retest agreement favored the OpenAI models.Conclusion Under standardized conditions, general-purpose multimodal LLMs, especially ChatGPT-5 and ChatGPT-4o, showed strong performance for psoriasis recognition and reasonable support for PGA scoring and treatment suggestions, supporting potential use by primary care physicians when dermatology specialist access is limited.| File | Dimensione | Formato | |
|---|---|---|---|
|
fmed-13-1791488.pdf
accesso aperto
Licenza:
Dominio pubblico
Dimensione
1.36 MB
Formato
Adobe PDF
|
1.36 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



