From the 1 of 19 linked papers with an AI index.
19 papers
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization
Peijie Chen, Zhuanling Zha, Zhipeng Nie +8
In current zero-shot text-to-speech systems, conventional semantic tokenizers are typically optimized using supervised automatic speech recognition or self-supervised learning obje…
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
Weijie Wu, Junbo Li, Lin Li +2
The paper introduces MMAC, a large benchmark of 5,638 audio clips designed to evaluate audio captioning models across multiple capability categories and evaluation dimensions, focu…
HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis
Wenhao Guan, Yifan Duan, Junxi Liu +6
Video dubbing is a cornerstone of multimedia content creation, aiming to synthesize synchronized acoustic sequences for visual streams. While Text-to-Speech (TTS) and Text-to-Audio…
SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations
Peijie Chen, Wenhao Guan, Weijie Wu +7
Zero-shot text-to-speech (TTS) relies on robust speech representations. However, current speech tokenizers face a fundamental trade-off: acoustic codecs preserve high-fidelity audi…
MeanFlowSE: one-step generative speech enhancement via conditional mean flow
Duojia Li, Shenghui Lu, Hongchen Pan +3
Multistep inference is a bottleneck for real-time generative speech enhancement because flow- and diffusion-based systems learn an instantaneous velocity field and therefore rely o…
Continual Audio Deepfake Detection via Universal Adversarial Perturbation
Wangjie Li, Lin Li, Qingyang Hong
The rapid advancement of speech synthesis and voice conversion technologies has raised significant security concerns in multimedia forensics. Although current detection models demo…