4 papers · 1 filter
Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition
Rishabh Jain, Aristeidis Papadopoulos, Zhaofeng Lin +1
Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the complexity of natural conversat…
Assessing True Generalisability of Audio-Visual Speech Recognisers
Zhaofeng Lin, Stavros Petridis, Maja Pantic +1
Current Audio-Visual Speech Recognition (AVSR) models achieve near-perfect performance on the standard LRS3 benchmark, raising concerns of adaptive overfitting. To systematically a…
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
Zhaofeng Lin, Naomi Harte
Audio-Visual Speech Recognition (AVSR) combines auditory and visual speech cues to enhance the accuracy and robustness of speech recognition systems. Recent advancements in AVSR ha…
Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation
Zhaofeng Lin, Tanvina Patel, Odette Scharenborg
Whispering is a distinct form of speech known for its soft, breathy, and hushed characteristics, often used for private communication. The acoustic characteristics of whispered spe…