activity
20242026
collaborators

6 papers

cs.SD2026

Leveraging large multimodal models for audio-video deepfake detection: a pilot study

Songjun Cao, Yuqi Li, Yunpeng Luo +2

Audio-visual deepfake detection (AVD) is increasingly important as modern generators can fabricate convincing speech and video. Most current multimodal detectors are small, task-sp…

cs.SD2025

FreeCodec: A disentangled neural speech codec with fewer tokens

Youqiang Zheng, Weiping Tu, Yueteng Kang +5

Neural speech codecs have gained great attention for their outstanding reconstruction with discrete token representations. It is a crucial component in generative tasks such as spe…

cs.SD2025

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis

Rui Niu, Weihao Wu, Jie Chen +2

Controllable speech synthesis aims to control the style of generated speech using reference input, which can be of various modalities. Existing face-based methods struggle with rob…

cs.SD2025

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt

Zhichao Wu, Yueteng Kang, Songjun Cao +3

Most existing Zero-Shot Text-To-Speech(ZS-TTS) systems generate the unseen speech based on single prompt, such as reference speech or text descriptions, which limits their flexibil…

cs.SD2025

DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models

Weihao wu, Zhiwei Lin, Yixuan Zhou +6

Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding o…

cs.SD2024

Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Xiong Wang, Yangze Li, Chaoyou Fu +5

Rapidly developing large language models (LLMs) have brought tremendous intelligent applications. Especially, the GPT-4o's excellent duplex speech interaction ability has brought i…