Publicly Available Large Language Models for Trichoscopy: A Head-to-Head Comparison with Dermatologists

    January 2026 in “ Diagnostics
    Basil Signer, Ali Mokhtari, Simone Cazzaniga, Flurin L. Brand, Gemma Caro, Pierre A. de Viragh, Kristine Heidemeyer, Aref Hosseini, Matilde Iorizzo, Alexandra Junge, Zora Martignoni, Pascal Reygagne, Bianca Maria Piraccini, Michela Starace, Antonia Reimer-Taschenbrecker, Charlotte Vogel, Dominik Obrist, S. Morteza Seyed Jafari
    This study compared the diagnostic accuracy of four publicly available large language models (LLMs) with that of 15 dermatologists, including residents, board-certified dermatologists, and trichology experts, in interpreting trichoscopic images. The dermatologists achieved a diagnostic accuracy of 58.1% for a single suspected diagnosis and 68.3% when including differential diagnoses, with experts outperforming others. In contrast, the LLMs had a significantly lower accuracy of 18.2% for a single diagnosis and 44.4% for including differential diagnoses, with Gemini 2.5 Flash performing best among the models. The agreement between AI models and dermatologists was slight to fair, indicating that LLMs currently underperform compared to human experts in trichology. The study concludes that LLMs need further development and specialized training to be reliable in assisting with trichological diagnoses.
    Discuss this study in the Community →

    Research cited in this study

    5 / 5 results