collaborators

5 papers

cs.SD2025

Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions

Wan Lin, Junhui Chen, Tianhao Wang +3

Modern speaker verification systems primarily rely on speaker embeddings, followed by verification based on cosine similarity between the embedding vectors of the enrollment and te…

cs.CV2025

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing

Zehua Liu, Xiaolou Li, Li Guo +2

Visual Speech Recognition (VSR) transcribes speech by analyzing lip movements. Recently, Large Language Models (LLMs) have been integrated into VSR systems, leading to notable perf…

cs.CV2025

CNVSRC 2024: The Second Chinese Continuous Visual Speech Recognition Challenge

Zehua Liu, Xiaolou Li, Chen Chen +2

This paper presents the second Chinese Continuous Visual Speech Recognition Challenge (CNVSRC 2024), which builds on CNVSRC 2023 to advance research in Chinese Large Vocabulary Con…

cs.SD2025

An Investigation on Speaker Augmentation for End-to-End Speaker Extraction

Zhenghai You, Zhenyu Zhou, Lantian Li +1

Target confusion, defined as occasional switching to non-target speakers, poses a key challenge for end-to-end speaker extraction (E2E-SE) systems. We argue that this problem is la…

cs.SD2024

AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition

Zehua Liu, Xiaolou Li, Chen Chen +3

Visual Speech Recognition (VSR) aims to recognize corresponding text by analyzing visual information from lip movements. Due to the high variability and weak information of lip mov…