6 papers
Leveraging large multimodal models for audio-video deepfake detection: a pilot study
Songjun Cao, Yuqi Li, Yunpeng Luo +2
Audio-visual deepfake detection (AVD) is increasingly important as modern generators can fabricate convincing speech and video. Most current multimodal detectors are small, task-sp…
FreeCodec: A disentangled neural speech codec with fewer tokens
Youqiang Zheng, Weiping Tu, Yueteng Kang +5
Neural speech codecs have gained great attention for their outstanding reconstruction with discrete token representations. It is a crucial component in generative tasks such as spe…
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis
Rui Niu, Weihao Wu, Jie Chen +2
Controllable speech synthesis aims to control the style of generated speech using reference input, which can be of various modalities. Existing face-based methods struggle with rob…
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
Zhichao Wu, Yueteng Kang, Songjun Cao +3
Most existing Zero-Shot Text-To-Speech(ZS-TTS) systems generate the unseen speech based on single prompt, such as reference speech or text descriptions, which limits their flexibil…
DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models
Weihao wu, Zhiwei Lin, Yixuan Zhou +6
Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding o…
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
Xiong Wang, Yangze Li, Chaoyou Fu +5
Rapidly developing large language models (LLMs) have brought tremendous intelligent applications. Especially, the GPT-4o's excellent duplex speech interaction ability has brought i…