Publicly Available Large Language Models for Trichoscopy: A Head-to-Head Comparison with Dermatologists
January 2026
in “
Diagnostics
”
This study compared the diagnostic accuracy of four publicly available large language models (LLMs) with that of 15 dermatologists, including residents, board-certified dermatologists, and trichology experts, in interpreting trichoscopic images. The dermatologists achieved a diagnostic accuracy of 58.1% for a single suspected diagnosis and 68.3% when including differential diagnoses, with experts outperforming others. In contrast, the LLMs had a significantly lower accuracy of 18.2% for a single diagnosis and 44.4% for including differential diagnoses, with Gemini 2.5 Flash performing best among the models. The agreement between AI models and dermatologists was slight to fair, indicating that LLMs currently underperform compared to human experts in trichology. The study concludes that LLMs need further development and specialized training to be reliable in assisting with trichological diagnoses.