collaborators

7 papers

cs.CV2026

GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining

Shentong Mo, Zehua Chen, Jun Zhu

Recent advances in video-audio (V-A) understanding and generation have increasingly relied on joint V-A embeddings, which serve as the foundation for tasks such as cross-modal retr…

cs.SD2025

Audio Super-Resolution with Latent Bridge Models

Chang Li, Zehua Chen, Liyuan Wang +1

Audio super-resolution (SR), i.e., upsampling the low-resolution (LR) waveform to the high-resolution (HR) version, has recently been explored with diffusion and bridge models, whi…

eess.AS2025

Deep Learning for Personalized Binaural Audio Reproduction

Xikun Lu, Yunda Chen, Zehua Chen +6

Personalized binaural audio reproduction is the basis of realistic spatial localization, sound externalization, and immersive listening, directly shaping user experience and listen…

cs.SD2025

FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio Generation

Yuxuan Jiang, Zehua Chen, Zeqian Ju +3

Text-to-audio (T2A) generation has achieved promising results with the recent advances in generative models. However, because of the limited quality and quantity of temporally-alig…

cs.LG2025

Versatile Cardiovascular Signal Generation with a Unified Diffusion Transformer

Zehua Chen, Yuyang Miao, Liyuan Wang +3

Cardiovascular signals such as photoplethysmography (PPG), electrocardiography (ECG), and blood pressure (BP) are inherently correlated and complementary, together reflecting the h…

cs.CV2025

DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap

Shentong Mo, Zehua Chen, Fan Bao +1

Recent works in cross-modal understanding and generation, notably through models like CLAP (Contrastive Language-Audio Pretraining) and CAVP (Contrastive Audio-Visual Pretraining),…