6 papers · 1 filter
DegDiT: Controllable Audio Generation with Dynamic Event Graph Guided Diffusion Transformer
Yisu Liu, Chenxing Li, Wanqian Zhang +6
Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and…
M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis
Xiaopeng Wang, Chunyu Qiang, Ruibo Fu +10
Non-autoregressive (NAR) text-to-speech synthesis relies on length alignment between text sequences and audio representations, constraining naturalness and expressiveness. Existing…
RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection
Ruibo Fu, Xiaopeng Wang, Zhengqi Wen +8
Existing methods for deepfake audio detection have demonstrated some effectiveness. However, they still face challenges in generalizing to new forgery techniques and evolving attac…
P2Mark: Plug-and-play Parameter-level Watermarking for Neural Speech Generation
Yong Ren, Jiangyan Yi, Tao Wang +7
Neural speech generation (NSG) has rapidly advanced as a key component of artificial intelligence-generated content, enabling the generation of high-quality, highly realistic speec…
Neural Codec Source Tracing: Toward Comprehensive Attribution in Open-Set Condition
Yuankun Xie, Xiaopeng Wang, Zhiyong Wang +7
Current research in audio deepfake detection is gradually transitioning from binary classification to multi-class tasks, referred as audio deepfake source tracing task. However, ex…
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
Hongming Guo, Ruibo Fu, Yizhong Geng +9
Text-to-audio (TTA) model is capable of generating diverse audio from textual prompts. However, most mainstream TTA models, which predominantly rely on Mel-spectrograms, still face…