collaborators

6 papers

cs.SD2026

SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents

Zeyu Xie, Chenxing Li, Qiao Jin +6

Recent audio generation models typically rely on Variational Autoencoders (VAEs) and perform generation within the VAE latent space. Although VAEs excel at compression and reconstr…

eess.AS2026

PhoenixCodec: Taming Neural Speech Coding for Extreme Low-Resource Scenarios

Zixiang Wan, Haoran Zhao, Guochang Zhang +3

This paper presents PhoenixCodec, a comprehensive neural speech coding and decoding framework designed for extremely low-resource conditions. The proposed system integrates an opti…

cs.SD2025

U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation

Xusheng Yang, Long Zhou, Wenfu Wang +6

We propose \textbf{U-Codec}, an \textbf{U}ltra low frame-rate neural speech \textbf{Codec} that achieves high-fidelity reconstruction and fast speech generation at an extremely low…

cs.SD2025

When Audio Generators Become Good Listeners: Generative Features for Understanding Tasks

Zeyu Xie, Chenxing Li, Xuenan Xu +6

This work pioneers the utilization of generative features in enhancing audio understanding. Unlike conventional discriminative features that directly optimize posterior and thus em…

cs.SD2025

FakeSound2: A Benchmark for Explainable and Generalizable Deepfake Sound Detection

Zeyu Xie, Yaoyun Zhang, Xuenan Xu +4

The rapid development of generative audio raises ethical and security concerns stemming from forged data, making deepfake sound detection an important safeguard against the malicio…

cs.SD2025

STAR: Speech-to-Audio Generation via Representation Learning

Zeyu Xie, Xuenan Xu, Yixuan Li +2

This work presents STAR, the first end-to-end speech-to-audio generation framework, designed to enhance efficiency and address error propagation inherent in cascaded systems. Unlik…