collaborators

7 papers

cs.SD2024

DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles

Jiaxuan Liu, Zhaoci Liu, Yajun Hu +3

Human speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffSty…

eess.AS2024

VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark

Yuke Lin, Ming Cheng, Fulin Zhang +3

In this paper, we provide a large audio-visual speaker recognition dataset, VoxBlink2, which includes approximately 10M utterances with videos from 110K+ speakers in the wild. This…

cs.CL2024

Exploring Energy-Based Models for Out-of-Distribution Detection in Dialect Identification

Yaqian Hao, Chenguang Hu, Yingying Gao +2

The diverse nature of dialects presents challenges for models trained on specific linguistic patterns, rendering them susceptible to errors when confronted with unseen or out-of-di…

eess.AS2024

GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model

Yingying Gao, Shilei Zhang, Chao Deng +1

Pre-trained speech language models such as HuBERT and WavLM leverage unlabeled speech data for self-supervised learning and offer powerful representations for numerous downstream t…

eess.AS2024

Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network

Yanan Chen, Zihao Cui, Yingying Gao +3

The expectation to deploy a universal neural network for speech enhancement, with the aim of improving noise robustness across diverse speech processing tasks, faces challenges due…

cs.LG2023

Cascaded Multi-task Adaptive Learning Based on Neural Architecture Search

Yingying Gao, Shilei Zhang, Zihao Cui +2

Cascading multiple pre-trained models is an effective way to compose an end-to-end system. However, fine-tuning the full cascaded model is parameter and memory inefficient and our…