5 papers
Audio Imitator: Controlling Timbre and Tempo in Video2Audio Synthesis with Audio Reference
Jiahui Zhao, Tianrui Wang, Chunyu Qiang +4
Video-to-audio generation has made significant progress in achieving semantic consistency and temporal alignment from silent videos. However, audio contains rich stylistic attribut…
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis
Bowen Li, Shaotong Guo, Zhen Wang +11
Building state-of-the-art text-to-speech (TTS) systems typically demands millions of hours of proprietary data and complex multi-stage architectures, creating substantial barriers…
InstructAudio: Unified speech and music generation with natural language instruction
Chunyu Qiang, Kang Yin, Xiaopeng Wang +8
Text-to-speech (TTS) and text-to-music (TTM) models face significant limitations in instruction-based control. TTS systems usually depend on reference audio for timbre, offer only…
Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation
Jun Wang, Xijuan Zeng, Chunyu Qiang +19
We propose Kling-Foley, a large-scale multimodal Video-to-Audio generation model that synthesizes high-quality audio synchronized with video content. In Kling-Foley, we introduce m…
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
Jiahui Zhao, Hao Shi, Chenrui Cui +5
Code-switching (CS) automatic speech recognition (ASR) faces challenges due to the language confusion resulting from accents, auditory similarity, and seamless language switches. A…