Authors: Emma Søndergaard Pedersen¹; Lukas Jürgensen¹; Søren Krogh Andersen²; Tobias Neher¹
¹Department of Clinical Research, University of Southern Denmark; ²Department of Psychology, University of Southern Denmark
Background: Successful speech communication in complex environments requires identifying a target signal amid several competing signals. Acoustic and linguistic similarity can cause peripheral and central auditory masking. Visual target speech can improve identification, whereas visual competing speech can cause interference. This study investigated how different auditory maskers and visual conditions affect audio-visual speech recognition.
Method: Speech recognition thresholds (SRT50s) were measured using the **Audio-Visual Danish Sentence Test (AV-DAST)** in 20 young adults with normal hearing and vision. Five scenarios were tested based on two auditory maskers (speech-shaped noise, SSN, and two-talker speech, TTS) and three visual conditions: audio-only (AO), visual target speech (AV1), and visual target plus competing speech (AV3). Visual benefit (AO-AV1) and disbenefit (AV1-AV3) were calculated, and test-retest reliability was evaluated after two weeks.
Results: In AO presentation, SSN and TTS yielded comparable SRT50s. With AV1, TTS resulted in lower SRT50s than SSN, showing a larger lip-reading benefit for TTS (3.9 dB) compared to SSN (2.2 dB). Visual competing speech caused a small overall disbenefit (mean: -0.5 dB) but showed large between-subject variability. Test-retest reliability was generally high across all measurements.
Conclusion: Visual target speech improves recognition more significantly in the presence of competing speech than in stationary noise. Young healthy adults vary markedly in their susceptibility to visual competing speech. These data provide a baseline for future research involving participants with sensory impairments.


