2 papers
eess.AS2025
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
Yi Wang, Oli Danyi Liu, Peter Bell
Human speech perception is multimodal. In natural speech, lip movements can precede corresponding voicing by a non-negligible gap of 100-300 ms, especially for specific consonants,…
cs.CL2025
Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
Michele Gubian, Ioana Krehan, Oli Liu +2
Analyses of self-supervised speech models have begun to reveal where and how they represent different types of information. However, almost all analyses have focused on English. He…