activity
20242026
collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2026

G-MaP-SE: Guided Speech Enhancement via GMM-Based Prior Matching

Yike Zhu, Ziqian Wang, Zikai Liu +5

Using speaker embeddings as conditioning can strengthen speech enhancement, but most methods either require clean enrollment audio or rely on embeddings extracted from noisy speech…

eess.AS2026

SVoice: Style-Aware Autoregressive Modeling with Enhanced Conditioning for Singing Style Conversion

Ziqian Wang, Xianjun Xia, Chuanzeng Huang +1

We present SVoice, the winning system of the Singing Voice Conversion Challenge (SVCC) 2025 for both the in-domain and zero-shot singing style conversion tracks. Built on the s…

eess.AS2025

LDCodec: A high quality neural audio codec with low-complexity decoder

Jiawei Jiang, Linping Xu, Dejun Zhang +3

Neural audio coding has been shown to outperform classical audio coding at extremely low bitrates. However, the practical application of neural audio codecs is still limited by the…

eess.AS2025

UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling

Ziqian Wang, Zikai Liu, Yike Zhu +6

Generative modeling has recently achieved remarkable success across image, video, and audio domains, demonstrating powerful capabilities for unified representation learning. Yet sp…

eess.AS2025

U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding

Ziqian Wang, Xianjun Xia, Xinfa Zhu +1

The text generation paradigm for audio tasks has opened new possibilities for unified audio understanding. However, existing models face significant challenges in achieving a compr…

eess.AS2024

BS-PLCNet 2: Two-stage Band-split Packet Loss Concealment Network with Intra-model Knowledge Distillation

Zihan Zhang, Xianjun Xia, Chuanzeng Huang +2

Audio packet loss is an inevitable problem in real-time speech communication. A band-split packet loss concealment network (BS-PLCNet) targeting full-band signals was recently prop…