collaborators

5 papers

eess.AS2026

SpeakerRPL v2: Robust Open-set Speaker Identification through Enhanced Few-shot Foundation Tuning and Model Fusion

Zhiyong Chen, Shuhang Wu, Yingjie Duan +2

This paper proposes an improved approach for open-set speaker identification based on pretrained speaker foundation models. Building upon the previous Speaker Reciprocal Points Lea…

cs.SD2025

RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval

Haoqin Sun, Jingguang Tian, Jiaming Zhou +8

The Contrastive Language-Audio Pretraining (CLAP) model has demonstrated excellent performance in general audio description-related tasks, such as audio retrieval. However, in the…

cs.SD2025

Learning Emotion-Invariant Speaker Representations for Speaker Verification

Jingguang Tian, Xinhui Hu, Xinkang Xu

In recent years, the rapid progress in speaker verification (SV) technology has been driven by the extraction of speaker representations based on deep learning. However, such repre…

cs.SD2025

Discrete Audio Representations for Automated Audio Captioning

Jingguang Tian, Haoqin Sun, Xinhui Hu +1

Discrete audio representations, termed audio tokens, are broadly categorized into semantic and acoustic tokens, typically generated through unsupervised tokenization of continuous…

eess.AS2025

A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model

Yang Xiang, Canan Huang, Desheng Hu +3

Most current speech enhancement (SE) methods recover clean speech from noisy inputs by directly estimating time-frequency masks or spectrums. However, these approaches often neglec…