Publicly Available Large Language Models for Trichoscopy: A Head-To-Head Comparison With Dermatologists

    January 2026 in “ Diagnostics
    Basil Signer, Ali Mokhtari, Simone Cazzaniga, Flurin L. Brand, Gemma Caro, Pierre A. de Viragh, Kristine Heidemeyer, Aref Hosseini, Matilde Iorizzo, Alexandra Junge, Zora Martignoni, Pascal Reygagne, Bianca Maria Piraccini, Michela Starace, Antonia Reimer-Taschenbrecker, Charlotte Vogel, Dominik Obrist, S. Morteza Seyed Jafari
    TLDR Large language models are less accurate than dermatologists in diagnosing trichoscopic images.
    This study evaluated the diagnostic accuracy of four publicly available large language models (LLMs) compared to 15 dermatologists in interpreting trichoscopic images for hair and scalp disorders. Dermatologists, including residents, board-certified dermatologists, and trichology experts, achieved higher diagnostic accuracy (58.1% for suspected diagnoses and 68.3% with differential diagnoses) than the AI models, which had an accuracy of 18.2% for suspected diagnoses and 44.4% with differential diagnoses. Among the AI models, Gemini performed best. The study concludes that LLMs currently underperform compared to human experts in trichology and require further development and specialized training to be reliable in routine care. Limitations include a small sample size and the use of pre-existing images, which may affect generalizability and accuracy.
    Discuss this study in the Community →

    Research cited in this study

    5 / 5 results