A Model of Continuous Speech Recognition Reveals the Role of Context in Human Speech Perception

Authors: Gasser Elbanna¹, Josh McDermott²

¹Harvard University
²Massachusetts Institute of Technology (MIT)

Background: Humans excel at transforming acoustic waveforms into meaningful linguistic representations, despite the inherent variability in speech signals. However, the underlying mechanisms that enable such robust perception remain unclear, as do their vulnerability to hearing loss. One bottleneck is the absence of models that replicate human performance and that could be used to probe for mechanistic hypotheses.

Methods: We developed PARROT, an artificial neural network model of continuous speech perception. The model combines a simulation of the human ear with convolutional and recurrent neural network modules (see Fig 1.a). These model stages map variable acoustic signals into linguistic representations (e.g., phonemes and characters). Trained on 7.5 million utterances, PARROT is the first phoneme-based recognition model at this scale. To evaluate human-model alignment, we designed a behavioral experiment in which participants transcribed non-words, enabling humans and models to be tested in the same way (see Fig 1.b). To study the role of contextual cues in human speech perception, we manipulated the model’s access to surrounding context.

Results: The experiment allowed us to compute a complete phoneme confusion matrix in humans for the first time, enabling a systematic comparison of human–model phoneme confusions. PARROT exhibited human-like patterns of phoneme confusions as well as accuracy (see Fig 2). Moreover, we found that models with access to both future and past context aligned more with human phonemic judgments than those using past or future alone. This result provides evidence that humans disambiguate speech sounds by integrating across a local time window that extends into the future (see Fig 3).

Conclusion: Overall, the results suggest that aspects of human-like speech perception emerge by optimizing for sub-lexical recognition from cochlear representations. Our work is a first step towards building biologically-plausible models that explain human speech encoding, and sets the stage for quantitative modeling of effects of hearing loss on speech perception.