3 papers
cs.CV2026
GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining
Shentong Mo, Zehua Chen, Jun Zhu
Recent advances in video-audio (V-A) understanding and generation have increasingly relied on joint V-A embeddings, which serve as the foundation for tasks such as cross-modal retr…
cs.SD2025
Audio Super-Resolution with Latent Bridge Models
Chang Li, Zehua Chen, Liyuan Wang +1
Audio super-resolution (SR), i.e., upsampling the low-resolution (LR) waveform to the high-resolution (HR) version, has recently been explored with diffusion and bridge models, whi…
cs.SD2025
FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio Generation
Yuxuan Jiang, Zehua Chen, Zeqian Ju +3
Text-to-audio (T2A) generation has achieved promising results with the recent advances in generative models. However, because of the limited quality and quantity of temporally-alig…