collaborators

9 papers

cs.SD2025

Investigating training objective for flow matching-based speech enhancement

Liusha Yang, Ziru Ge, Gui Zhang +2

Speech enhancement(SE) aims to recover clean speech from noisy recordings. Although generative approaches such as score matching and Schrodinger bridge have shown strong effectiven…

cs.SD2025

AnyAccomp: Generalizable Accompaniment Generation via Quantized Melodic Bottleneck

Junan Zhang, Yunjia Zhang, Xueyao Zhang +1

Singing Accompaniment Generation (SAG) is the process of generating instrumental music for a given clean vocal input. However, existing SAG techniques use source-separated vocals a…

cs.SD2025

The CCF AATC 2025 Speech Restoration Challenge: A Retrospective

Junan Zhang, Mengyao Zhu, Xin Xu +3

Real-world speech communication is rarely affected by a single type of degradation. Instead, it suffers from a complex interplay of acoustic interference, codec compression, and, i…

cs.SD2025

TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling

Yuancheng Wang, Dekun Chen, Xueyao Zhang +3

Speech tokenizers serve as foundational components for speech language models, yet current designs exhibit several limitations, including: 1) dependence on multi-layer residual vec…

cs.SD2025

Multi-Metric Preference Alignment for Generative Speech Restoration

Junan Zhang, Xueyao Zhang, Jing Yang +3

Recent generative models have significantly advanced speech restoration tasks, yet their training objectives often misalign with human perceptual preferences, resulting in suboptim…

cs.SD2025

SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset

Yicheng Gu, Chaoren Wang, Junan Zhang +4

The lack of a publicly-available large-scale and diverse dataset has long been a significant bottleneck for singing voice applications like Singing Voice Synthesis (SVS) and Singin…