Background: Non-infectious inflammatory skin diseases are a common reason for dermatological consultation, and their overlapping clinical features make them diagnostically challenging. Patients increasingly use conversational artificial intelligence (AI) chatbots to interpret their symptoms before seeking medical care. However, these tools have rarely been evaluated on the everyday, non-technical language that patients use to describe a skin problem. Objectives: This study aimed to evaluate and compare the diagnostic performance of four conversational AI platforms (three general-purpose chatbots and one clinically oriented platform) against a dermatologist-confirmed reference diagnosis, using patient-reported descriptions (PRDs) of non-infectious inflammatory skin diseases, and to describe how each platform structured its responses. Methods: In this comparative study, the same 67 PRDs, written in the patients' own words, were submitted once to ChatGPT, Gemini, Claude, and OpenEvidence; each description closed with a standardised lay question ("What could I have?"). Responses were scored on two levels: whether the first proposed diagnosis matched the reference (Top-1) and whether the correct diagnosis appeared among the first four hypotheses (Top-4). Sensitivity and specificity were calculated for the three conditions most represented in the dataset (psoriasis, hidradenitis suppurativa, and acne). Additional features of the responses were also recorded descriptively, such as their communicative profile and whether referral to a physician was suggested. Results: Top-1 accuracy was moderate and differed between platforms, ranging from 56.7% (ChatGPT) to 71.6% (Gemini). At Top-4, accuracy rose markedly and converged across platforms (86.6%-89.6%), indicating that the correct diagnosis was often present in the broader differential but not always ranked first. Across conditions, sensitivity increased markedly from Top-1 to Top-4, whereas specificity remained consistently low and fell below 17% at Top-4. This low specificity reflected the tendency of the platforms to propose broad differential lists, generating frequent false positives. The clinically oriented platform did not outperform the general-purpose chatbots (Top-1 64.2% vs 64.7%; Top-4 86.6% vs 87.6%); the main differences between tools concerned their communicative profile. Conclusions: Conversational AI platforms recognised the correct condition within a broad differential in most cases but frequently failed to rank it first. They may therefore support preliminary orientation and patient education but should not be regarded as autonomous diagnostic tools. Dermatological assessment, including visual examination, and clinical supervision remain essential to safe and accurate diagnosis.

ARTIFICIAL INTELLIGENCE FOR THE DIAGNOSIS OF INFLAMMATORY SKIN DISEASES

MADEO, MARTINA FATIMA
2025/2026

Abstract

Background: Non-infectious inflammatory skin diseases are a common reason for dermatological consultation, and their overlapping clinical features make them diagnostically challenging. Patients increasingly use conversational artificial intelligence (AI) chatbots to interpret their symptoms before seeking medical care. However, these tools have rarely been evaluated on the everyday, non-technical language that patients use to describe a skin problem. Objectives: This study aimed to evaluate and compare the diagnostic performance of four conversational AI platforms (three general-purpose chatbots and one clinically oriented platform) against a dermatologist-confirmed reference diagnosis, using patient-reported descriptions (PRDs) of non-infectious inflammatory skin diseases, and to describe how each platform structured its responses. Methods: In this comparative study, the same 67 PRDs, written in the patients' own words, were submitted once to ChatGPT, Gemini, Claude, and OpenEvidence; each description closed with a standardised lay question ("What could I have?"). Responses were scored on two levels: whether the first proposed diagnosis matched the reference (Top-1) and whether the correct diagnosis appeared among the first four hypotheses (Top-4). Sensitivity and specificity were calculated for the three conditions most represented in the dataset (psoriasis, hidradenitis suppurativa, and acne). Additional features of the responses were also recorded descriptively, such as their communicative profile and whether referral to a physician was suggested. Results: Top-1 accuracy was moderate and differed between platforms, ranging from 56.7% (ChatGPT) to 71.6% (Gemini). At Top-4, accuracy rose markedly and converged across platforms (86.6%-89.6%), indicating that the correct diagnosis was often present in the broader differential but not always ranked first. Across conditions, sensitivity increased markedly from Top-1 to Top-4, whereas specificity remained consistently low and fell below 17% at Top-4. This low specificity reflected the tendency of the platforms to propose broad differential lists, generating frequent false positives. The clinically oriented platform did not outperform the general-purpose chatbots (Top-1 64.2% vs 64.7%; Top-4 86.6% vs 87.6%); the main differences between tools concerned their communicative profile. Conclusions: Conversational AI platforms recognised the correct condition within a broad differential in most cases but frequently failed to rank it first. They may therefore support preliminary orientation and patient education but should not be regarded as autonomous diagnostic tools. Dermatological assessment, including visual examination, and clinical supervision remain essential to safe and accurate diagnosis.
2025
AI
Dermatology
Dermatoses
Patient symptoms
Diagnostic accuracy
File in questo prodotto:
File Dimensione Formato  
Madeo.MartinaFatima.pdf

Accesso riservato

Dimensione 1.67 MB
Formato Adobe PDF
1.67 MB Adobe PDF

I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14251/6772