collaborators
Showing cs.SDShow all

7 papers · 1 filter

cs.SD2025

Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions

Wan Lin, Junhui Chen, Tianhao Wang +3

Modern speaker verification systems primarily rely on speaker embeddings, followed by verification based on cosine similarity between the embedding vectors of the enrollment and te…

cs.SD2025

An Investigation on Speaker Augmentation for End-to-End Speaker Extraction

Zhenghai You, Zhenyu Zhou, Lantian Li +1

Target confusion, defined as occasional switching to non-target speakers, poses a key challenge for end-to-end speaker extraction (E2E-SE) systems. We argue that this problem is la…

cs.SD2024

AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition

Zehua Liu, Xiaolou Li, Chen Chen +3

Visual Speech Recognition (VSR) aims to recognize corresponding text by analyzing visual information from lip movements. Due to the high variability and weak information of lip mov…

cs.SD2024

Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective

Chen Chen, Xiaolou Li, Zehua Liu +2

In the field of spoken language processing, audio-visual speech processing is receiving increasing research attention. Key components of this research include tasks such as lip rea…

cs.SD2024

Zero-Shot Fake Video Detection by Audio-Visual Consistency

Xiaolou Li, Zehua Liu, Chen Chen +3

Recent studies have advocated the detection of fake videos as a one-class detection task, predicated on the hypothesis that the consistency between audio and visual modalities of g…

cs.SD2024

SE/BN Adapter: Parametric Efficient Domain Adaptation for Speaker Recognition

Tianhao Wang, Lantian Li, Dong Wang

Deploying a well-optimized pre-trained speaker recognition model in a new domain often leads to a significant decline in performance. While fine-tuning is a commonly employed solut…