Authors: Wiktor Jedrzejczak¹
¹Institute of Physiology and Pathology of Hearing
Background: Chatbots powered by large language models (LLMs) have recently emerged as prominent sources of information. However, their ability to propagate misinformation as well as information, particularly in specialized fields like audiology, remains underexplored. This study aimed to evaluate the accuracy of six popular chatbots – ChatGPT, Gemini, Claude, DeepSeek, Grok and Mistral – in response to questions related to unproven methods in audiological care.
Methods: A set of 50 questions was developed based on common inquiries from patient-clinician conversations. These questions were posed to chatbots. The chatbots were tested 10 times to account for variability of responses, giving 3,000 total responses. Responses were compared with key based on averaged opinion by 11 professional experts for accuracy and alignment with evidence-based knowledge. The consistency of responses was evaluated by Cohen’s Kappa.
Results: For the majority of questions, most chatbot responses were deemed accurate. Grok consistently performed best, achieving 100% alignment with expert opinions. Deepseek exhibited the lowest accuracy, scoring 95.8%. Mistral exhibited lowest consistency, scoring 0.96.
Conclusions: Although evaluated chatbots generally refrained from endorsing scientifically unsupported methods, certain inconsistencies could still facilitate misinformation. Outstandingly, Grok provided accurate and reliable responses, underscoring its potential utility in clinical and educational settings.

