Machine learning models of hearing demonstrate the limits of attentional selection of speech heard through cochlear implants

Authors: Annesya Banerjee¹, Ian Griffith¹, Josh McDermott²

¹Harvard University
²Massachusetts Institute of Technology (MIT)

Background: Humans with normal hearing abilities are able to attend to target sources in the presence of concurrent sounds, allowing them to communicate in noisy environments. Such abilities are limited for individuals with hearing loss and users of cochlear implants. Attentional deficits could reflect degraded peripheral information, for instance if attentional cues are not encoded with sufficient fidelity. Alternatively, deficits could result from suboptimal decoding of the altered peripheral representations that follow hearing loss and cochlear implantation (as if the central auditory system cannot fully adapt to the altered periphery). To study this issue, we optimized artificial neural network models to recognize speech from a cued talker in multi-talker settings, and asked whether the models could perform the task using simulated cochlear input stimulation.

Method: We optimized deep neural networks to report words spoken by a cued talker in a multi-source mixture. Models were trained using simulated binaural auditory nerve input obtained from either a normal cochlea, a cochlea with degraded temporal coding (simulated via lowering the nerve phase-locking cutoff to 50 Hz), or a simulated cochlear implant. Attentional selection was enabled by stimulus-computable feature-based gains, implemented with learnable logistic functions operating on the time-averaged model activations of a cued talker. Gains could be high for features of the cue, and low for uncued features, as determined by parameters optimized to maximize task performance.

Results: Models with normal nerve input successfully learned to use both spatial and vocal timbre cues to solve the word recognition task. In the presence of competing talkers, these models correctly reported the words of the cued talker and ignored the distractor talker(s), similar to humans with normal hearing abilities. Models with degraded temporal coding performed worse than the normal hearing model, but showed some benefit of target-distractor spatial separation and sex difference. Models with simulated cochlear implant stimulation performed notably worse, showing only modest benefits from target-distractor sex differences, and showing spatial benefits only for very large spatial separation.

Conclusion: Our results suggest that auditory attention deficits in cochlear implant users reflect limitations of peripheral information available from current electrical stimulation strategies.