7 citations · 7 across the 7 of their papers we have counts for
7 papers
Discrete vs. Continuous: A Comprehensive Study of Unified Audio Understanding in LALMs
Jing Peng, Zichao Nie, Zhisheng Zhang +2
Large Audio Language Models (LALMs) utilize either continuous features or discrete tokens, yet the optimal representation paradigm for general audio understanding remains debated.…
Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions
Haiyun Li, Shuhai Peng, Zhisheng Zhang +4
Audio watermarking aims to embed identifiable information into audio while remaining imperceptible. Existing methods adopt high-fidelity, low-energy designs to preserve perceptual…
VoxCPM2 Technical Report
Yixuan Zhou, Guoyang Zeng, Xin Liu +15
We present VoxCPM2, a https://info.arxiv.org/help/prep#abstractsfully open-source multilingual and controllable speech generation foundation model that extends the hierarchical dif…
LoSATok: Low-Dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation
Zhisheng Zhang, Xiang Li, Yixuan Zhou +3
Audio tokenizers are fundamental to unifying audio understanding and generation. Understanding requires high-level semantics, while generation demands semantic and acoustic details…
SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis
Zhisheng Zhang, Derui Wang, Qianyi Yang +6
Speech synthesis technology has brought great convenience, while the widespread usage of realistic deepfake audio has triggered hazards. Malicious adversaries may unauthorizedly co…
Mitigating Unauthorized Speech Synthesis for Voice Protection
Zhisheng Zhang, Qianyi Yang, Derui Wang +4
With just a few speech samples, it is possible to perfectly replicate a speaker's voice in recent years, while malicious voice exploitation (e.g., telecom fraud for illegal financi…