- Research Article
- 10.1080/25742442.2025.2593218
Integration of Contrastive Prosody and Segmental Information During Spoken Word Recognition
- Nov 27, 2025
- Auditory Perception & Cognition
- Andrés Buxó-Lugo
ABSTRACT Listeners use many sources of information to make sense of speech sounds and the words these sounds are communicating. Segmental cues, such as power fluctuations across frequency bands, allow to distinguish between speech sounds. Prosody – the rhythmic and intonational aspects of speech – provides high-level information for speech comprehension. The present study investigates how listeners combine segmental and prosodic information during spoken word recognition. In a modified two alternative forced-choice task, participants heard two words and provided judgments of what word they heard for each. Critical trials consisted of/b/-/p/minimal pairs (e.g., bin-pin) in which voice onset time and aspiration of the second word were manipulated along a 9-step continuum. Critically, the intonational contour of the second word was also manipulated so that it either had contrastive or neutral intonation. The acoustic details of the first segment of the second word were also manipulated. Results show that listeners balance segmental and prosodic information in complex ways, where the effect of contrastive intonation varies depending on the ambiguity of the/b/-/p/onsets based on the VOT and aspiration manipulation, as well as the identity of the word that the current word is being contrasted from. Results were compared to predictions from three computational models to test which cue integration process yielded more human-like responses. The model that yielded the best fit proposed that the effect of prosody was weighed continuously as a function of the ambiguity of the segmental cues, suggesting a highly interactive process of integration.
Read more