Automated scoring of speech audiometry for cochlear implant listeners using artificial intelligence and flexible grading methods

Authors: Erdem Baha Topbas¹, Tobias Goehring¹

¹Deep Hearing Lab, MRC Cognition and Brain Sciences Unit

Background: Web-based remote speech audiometry enables larger-scale and more frequent data collection. Traditional scoring methods are costly to handle this quantity of data and incompatible with remote adaptive testing in real-time. This project explores using general-purpose automatic speech recognition (ASR) models to transcribe response recordings to open-set speech audiometry tests from cochlear implant (CI) users.

Method: CI users (n=15) listened to and repeated sentences from the BKB corpus in two test locations (in the lab and remotely using personal devices at home) in three listening conditions (clean speech, in babble noise, after DNN noise reduction). Human scorers unfamiliar with the corpus (n=4) and state-of-the-art ASR models (n=7) transcribed responses. Transcripts were scored using tight and loose scoring rules. Traditional keyword scoring by human scorers familiar with the corpus (n=4) served as benchmark.

Results: Benchmark accuracy averaged keyword scores of 60.7%. Unfamiliar human scorers averaged 52.7%, and the best ASR model produced 49.3%. ASR scores were closer to human scores for the lab-based testing. The ASR score was closer to the benchmark for the noisier stimuli than the clean ones. Cohen’s d effect sizes and intraclass correlation coefficients of subject rankings showed a high correlation between ASR and benchmark results. Loose scoring led to improved accuracy, closer to human scores, across conditions. Scorers were affected differently by loose scoring for effect size and ICC results. Transcription time was reduced by 75%.

Conclusion: The ASR models approached benchmark scores and preserved test outcomes, with similar performance as unfamiliar human scorers. They facilitate fast and cost-effective scoring for large-scale speech audiometry, especially in academic studies or clinical testing for within-subject effects. However, differences in absolute scores to the benchmark may limit their application for clinical assessments.