2 papers
cs.CV2026
The Lipreading Gap: Do VSR Models Perceive Visual Speech Like Human Lipreaders?
Rishabh Jain, Naomi Harte
Visual speech recognition (VSR) models now surpass human lipreaders on benchmarks, but do such gains establish human-like visual speech perception? To explore this, we compare thre…
eess.AS2026
VisG AV-HuBERT: Viseme-Guided AV-HuBERT
Aristeidis Papadopoulos, Rishabh Jain, Naomi Harte
Audio-Visual Speech Recognition (AVSR) systems nowadays integrate Large Language Model (LLM) decoders with transformer-based encoders, achieving state-of-the-art results. However,…