collaborators

7 papers

eess.AS2026

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling

Guanrou Yang, Tian Tan, Qian Chen +12

Integrating speech understanding and generation is a pivotal step toward building unified speech models. However, the different representations required for these two tasks current…

cs.IR2025

CART: A Generative Cross-Modal Retrieval Framework with Coarse-To-Fine Semantic Modeling

Minghui Fang, Shengpeng Ji, Jialong Zuo +9

Cross-modal retrieval aims to search for instances, which are semantically related to the query through the interaction of different modal data. Traditional solutions utilize a sin…

eess.AS2025

Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding

Rui-Chen Zheng, Wenrui Liu, Hui-Peng Du +6

Existing speech tokenizers typically assign a fixed number of tokens per second, regardless of the varying information density or temporal fluctuations in the speech signal. This u…

eess.AS2025

EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting

Guanrou Yang, Chen Yang, Qian Chen +12

Human speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made h…

eess.AS2025

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation

Wenrui Liu, Qian Chen, Wen Wang +11

Neural audio codecs, used as speech tokenizers, have demonstrated remarkable potential in the field of speech generation. However, to ensure high-fidelity audio reconstruction, neu…

cs.SD2025

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

Jialong Zuo, Shengpeng Ji, Minghui Fang +8

This paper introduces PFlow-VC, a conditional flow matching voice conversion model that leverages fine-grained discrete pitch tokens and target speaker prompt information for expre…