3 papers
cs.SD2025
Semantic Codebooks as Effective Priors for Neural Speech Compression
Liuyang Bai, Weiyi Lu, Li Guo
Speech codecs are traditionally optimized for waveform fidelity, allocating bits to preserve acoustic detail even when much of it can be inferred from linguistic structure. This le…
cs.CV2025
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
Zehua Liu, Xiaolou Li, Li Guo +2
Visual Speech Recognition (VSR) transcribes speech by analyzing lip movements. Recently, Large Language Models (LLMs) have been integrated into VSR systems, leading to notable perf…
cs.SD2024
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
Zehua Liu, Xiaolou Li, Chen Chen +3
Visual Speech Recognition (VSR) aims to recognize corresponding text by analyzing visual information from lip movements. Due to the high variability and weak information of lip mov…