How hard is it to read a chatbot response about hearing health

🤖💬 How Readable Are AI-Generated Hearing Health Responses?

Hector Gabriel Corrale de Matos
Jan-Willem A. Wasmann
Lilian Cássia Bornia Jacob

Short summary: Chatbots offer highly accessible hearing-health advice, however their readability in different languages remains unclear. This study examines whether incorporating Wikipedia-derived audiology excerpts alters the textual complexity of AI-generated responses in English and Brazilian Portuguese, evaluated with standard readability indices.

 

 

📢 Help us out! Please support this multilingual readability study by filling out the short survey:

🕒 It only takes a few minutes. Thank you very much for your help!

Your input is highly relevant. We will present the final results here and in a publication.

We are evaluating how clearly AI generated texts communicate WHO recommendations on hearing aid service delivery in low and middle income countries (LMICs). Your clinical expertise is crucial.

Ethical Consideration: This project does not meet the regulatory definitions of human subjects research. In line with Brazilian National Health Council Resolution No. 510/2016, Title 45 CFR Part 46 (Common Rule), and the World Medical Association Declaration of Helsinki (2013, amended 2024), institutional review-board approval is not required. Expert evaluators act independently; no identifiable personal data are collected; participation is anonymous, voluntary, and entails no greater than minimal risk.

📝 What to do

  • Read the one-page Required Reading
  • Rate 16 short AI-generated answers.
  • ≈15 min, fully anonymous, and no personal data collected.

📄 Reference documents

Thank you for advancing responsible, evidence-based AI in audiology! Your contribution is extremely valuable!

Study team: Hector G. C. de Matos (University of São Paulo), Jan-Willem Wasmann (Radboudumc), Lilian C. B. Jacob (University of São Paulo).

🔗 Choose a survey link (select one or more at random):

Please select randomly one or more links and respond carefully. You don’t need to identify yourself at any point; your ratings join an independent panel-of-judges system for unbiased evaluation.

🌍 English 
🌍 Brazilian Portuguese
This coming 13th of May

📌 More information:

📖 Background

Large language models (LLMs) are playing an increasing role in how people learn about audiology and hearing health. Alongside OpenAI’s ChatGPT, a new generation — including DeepSeek from China and MaritacaAI from Brazil — now offer regional users direct access to information. The usefulness of these tools depends in part on how easily everyday readers can understand the language they produce.Readability metrics including Flesch Reading Ease, Gunning Fog, SMOG, and Dale–Chall measure how hard a text is to read. They usually rely on features like sentence length, number of syllables, or how common the words are—assuming human writers and a steady writing style. LLMs often mix technical terms with everyday language, shift tone suddenly, and pack complex ideas into short clauses to mislabel real difficulty. Studies show that Generative Pre-trained Transformer self-ratings correlate with expert opinion, still their accuracy drifts when prompts, genres, or audiences change, confirming that comprehension is context-bound.

Retrieval-augmented generation adds further complexity. By injecting passages from repositories such as Wikipedia, augmented generation grounds answers in collaboratively vetted knowledge, reducing hallucination and sometimes easing and sometimes hampering comprehension. Wikipedia’s style can affect readability, and its content varies across languages in depth and clarity. Syllable-based metrics fail for character-based scripts like Chinese and endeavor with agglutinative languages; no universal index exists.

These developments imply more than technical enhancements; they hold operative implications for determining whether AI-mediated hearing health communication serves to mitigate or exacerbate existing health inequities. These metrics may thus serve as a proxy for the equitable deployment of LLMs, particularly when machine-generated text is grounded in the collaboratively produced knowledge of Wikipedia.

💻Code examples on readability metrics and online tools
Python, Julia and R.
📚 References

Agency for Healthcare Research and Quality. (2024). Personal health literacy [Primer]. https://psnet.ahrq.gov/primer/personal-health-literacy

Amstad, T. (1978). Wie verständlich sind unsere Zeitungen? [Doctoral dissertation, University of Zurich].

Bansal, S., & Contributors. (2025). textstat (Version 0.7.5) [Computer software]. Python Package Index. https://pypi.org/project/textstat/

Barrio-Cantalejo, I., Simón-Lorda, P., & Puerta-Melguizo, M. C. (2008). Validación de la escala INFLESZ para evaluar la legibilidad de textos dirigidos al paciente. Anales del Sistema Sanitario de Navarra, 31(2), 135–152.

Benoit, K., Watanabe, K., Wang, H., Müller, S., & Obeng, A. (2024). quanteda: An R package for quantitative text analysis (Version 2.1.2) [Computer software]. https://quanteda.io/reference/textstat_readability.html

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., … Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.

Centers for Disease Control and Prevention. (2010). Simply Put: A guide for creating easy-to-understand materials (3rd ed.). U.S. Department of Health and Human Services. https://stacks.cdc.gov/view/cdc/11938

Centers for Disease Control and Prevention. (2024a). Plain language materials & resources. https://www.cdc.gov/health-literacy/php/develop-materials/plain-language.html

Centers for Disease Control and Prevention. (2024b). Plain writing at CDC. https://www.cdc.gov/other/plainwriting.html

Coleman CK, Muñoz K, Ong CW, Butcher GM, Nelson L, Twohig M. Opportunities for Audiologists to Use Patient-Centered Communication during Hearing Device Monitoring Encounters. Semin Hear. 2018 Feb;39(1):32-43. doi: 10.1055/s-0037-1613703. Epub 2018 Feb 7. PMID: 29422711; PMCID: PMC5802989.

Dale, E., & Chall, J. S. (1948). A formula for predicting readability. Educational Research Bulletin, 27(1), 11–20.

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (pp. 4171–4186). https://doi.org/10.18653/v1/N19-1423

Eltorai, A. E. M., Ghanian, S., Adams, C. A., Born, C. T., & Daniels, A. H. (2014). Readability of patient education materials on the American Association for Surgery of Trauma website. Archives of Trauma Research, 3(2), e18161. https://doi.org/10.5812/atr.18161

Fernández-Huerta, J. (1959). Medidas sencillas de lecturabilidad. Consigna, 214, 29–32.

Flesch, R. (1948). A new readability yardstick. Journal of Applied Psychology, 32(3), 221–233. https://doi.org/10.1037/h0057532

Gunning, R. (1952). The technique of clear writing. McGraw-Hill.

JuliaText. (2024). TextAnalysis.jl (Version 0.8.2) [Computer software]. https://github.com/JuliaText/TextAnalysis.jl

Kincaid, J. P., Fishburne, R. P., Rogers, R. L., & Chissom, B. S. (1975). Derivation of new readability formulas for Navy enlisted personnel (Research Branch Report 8-75). Naval Technical Training Command.

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., … Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474.

Manchaiah V, Bellon-Harn ML, Dockens AL, Azios JH, Harn WE. Communication between Audiologist, Patient, and Patient’s Family Members during Initial Audiology Consultation and Rehabilitation Planning Sessions: A Descriptive Review. J Am Acad Audiol. 2019 Oct;30(9):810-819. doi: 10.3766/jaaa.18032. Epub 2018 Nov 29. PMID: 30541658.

Martins, T. B. F., Ghiraldelo, C. M., Nunes, M. G. V., & Oliveira, O. N. (1996). Readability formulas applied to textbooks in Brazilian Portuguese. Notas do ICMSC – Série Computação, 28, 1–24.

McLaughlin, G. H. (1969). SMOG grading: A new readability formula. Journal of Reading, 12(8), 639–646.

Morata, T. C., Zucki, F., Arrigo, A. J., Cruz, P. C., Gong, W., Matos, H. G. C., Montilha, A. A. P., Peschanski, J. A., Cardoso, M. J., Lacerda, A. B. M., Berberian, A. P., Araujo, E. S., Luders, D., Duarte, J. L., Jacob, R. T. S., Chadha, S., Mietchen, D., Rasberry, L., Alvarenga, K. F., & Jacob, L. C. B. (2024). Strategies for crowdsourcing hearing health information: a comparative study of educational programs and volunteer-based campaigns on Wikimedia. BMC public health24(1), 2646. https://doi.org/10.1186/s12889-024-20105-8

Montilha, A. A. P., Morata, T. C., Flor, D. Á., Machado, M. A. A. M., Menegon, F. A., & Zucki, F. (2023). The Promotion of Hearing Health through Wikipedia Campaigns: Article Quality and Reach Assessment. Healthcare (Basel, Switzerland)11(11), 1572. https://doi.org/10.3390/healthcare11111572

National Institutes of Health. (2003). Clear & simple: Developing effective print materials for low-literacy readers (NIH Pub. No. 03-0865). U.S. Department of Health and Human Services.

Stacey JE, Atkin C, Roberts KL, Henshaw H, Allen HA, Badham SP. Memory for health information: Influences of age, hearing aids, and multisensory presentation. Q J Exp Psychol (Hove). 2024 Nov 22:17470218241295722. doi: 10.1177/17470218241295722. Epub ahead of print. PMID: 39439103.

Szigriszt-Pazos, F. (1993). Sistemas predictivos de legibilidad del mensaje escrito: Fórmula de perspicuidad [Doctoral dissertation, Universidad Complutense de Madrid].

Sung, Y.-T., Chang, T.-H., Lin, W.-C., Hsieh, K.-S., & Chang, K.-E. (2016). CRIE: An automated analyzer for Chinese texts. Behavior Research Methods, 48(4), 1238–1251. https://doi.org/10.3758/s13428-015-0649-1

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008.