2 citations · 2 across the 5 of their papers we have counts for
5 papers
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
Wenrui Liu, Qian Chen, Wen Wang +11
Neural audio codecs, used as speech tokenizers, have demonstrated remarkable potential in the field of speech generation. However, to ensure high-fidelity audio reconstruction, neu…
Continual Cross-Modal Generalization
Yan Xia, Hai Huang, Minghui Fang +1
Cross-modal generalization aims to learn a shared discrete representation space from multimodal pairs, enabling knowledge transfer across unannotated modalities. However, achieving…
Speech Watermarking with Discrete Intermediate Representations
Shengpeng Ji, Ziyue Jiang, Jialong Zuo +4
Speech watermarking techniques can proactively mitigate the potential harmful consequences of instant voice cloning techniques. These techniques involve the insertion of signals in…
WavChat: A Survey of Spoken Dialogue Models
Shengpeng Ji, Yifu Chen, Minghui Fang +16
Recent advancements in spoken dialogue models, exemplified by systems like GPT-4o, have captured significant attention in the speech domain. Compared to traditional three-tier casc…
MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and Correspondence
Fuming You, Minghui Fang, Li Tang +3
Motion-to-music and music-to-motion have been studied separately, each attracting substantial research interest within their respective domains. The interaction between human motio…