activity
20242026
collaborators

9 papers

cs.LG2026

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1

With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…

cs.SD2025

Training Text-to-Speech Model with Purely Synthetic Data: Feasibility, Sensitivity, and Generalization Capability

Tingxiao Zhou, Leying Zhang, Zhengyang Chen +1

The potential of synthetic data in text-to-speech (TTS) model training has gained increasing attention, yet its rationality and effectiveness require systematic validation. In this…

cs.SD2025

Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction

Leying Zhang, Wangyou Zhang, Zhengyang Chen +1

The acoustic background plays a crucial role in natural conversation. It provides context and helps listeners understand the environment, but a strong background makes it difficult…

cs.SD2024

Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning

Bing Han, Wen Huang, Zhengyang Chen +7

The goal of the acoustic scene classification (ASC) task is to classify recordings into one of the predefined acoustic scene classes. However, in real-world scenarios, ASC systems…

eess.AS2024

Prototype and Instance Contrastive Learning for Unsupervised Domain Adaptation in Speaker Verification

Wen Huang, Bing Han, Zhengyang Chen +2

Speaker verification system trained on one domain usually suffers performance degradation when applied to another domain. To address this challenge, researchers commonly use featur…

cs.SD2024

Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching

Zhengyang Chen, Bing Han, Shuai Wang +2

Speaker diarization is typically considered a discriminative task, using discriminative approaches to produce fixed diarization results. In this paper, we explore the use of neural…