activity
20242026
collaborators

9 papers

cs.SD2026

DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer

Yisu Liu, Chenxing Li, Wanqian Zhang +6

Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and…

cs.SD2025

M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis

Xiaopeng Wang, Chunyu Qiang, Ruibo Fu +10

Non-autoregressive (NAR) text-to-speech synthesis relies on length alignment between text sequences and audio representations, constraining naturalness and expressiveness. Existing…

cs.SD2025

RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection

Ruibo Fu, Xiaopeng Wang, Zhengqi Wen +8

Existing methods for deepfake audio detection have demonstrated some effectiveness. However, they still face challenges in generalizing to new forgery techniques and evolving attac…

cs.MM2025

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model

Yong Ren, Chenxing Li, Le Xu +7

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remain…

cs.SD2025

P2Mark: Plug-and-play Parameter-level Watermarking for Neural Speech Generation

Yong Ren, Jiangyan Yi, Tao Wang +7

Neural speech generation (NSG) has rapidly advanced as a key component of artificial intelligence-generated content, enabling the generation of high-quality, highly realistic speec…

cs.MM2025

Deconfounded Reasoning for Multimodal Fake News Detection via Causal Intervention

Moyang Liu, Kaiying Yan, Yukun Liu +4

The rapid growth of social media has led to the widespread dissemination of fake news across multiple content forms, including text, images, audio, and video. Traditional unimodal…