Showing cs.SDShow all
3 papers · 1 filter
cs.SD2024
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
Zehua Liu, Xiaolou Li, Chen Chen +3
Visual Speech Recognition (VSR) aims to recognize corresponding text by analyzing visual information from lip movements. Due to the high variability and weak information of lip mov…
cs.SD2024
Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective
Chen Chen, Xiaolou Li, Zehua Liu +2
In the field of spoken language processing, audio-visual speech processing is receiving increasing research attention. Key components of this research include tasks such as lip rea…
cs.SD2024
Zero-Shot Fake Video Detection by Audio-Visual Consistency
Xiaolou Li, Zehua Liu, Chen Chen +3
Recent studies have advocated the detection of fake videos as a one-class detection task, predicated on the hypothesis that the consistency between audio and visual modalities of g…