5 papers
Cinematic Audio Source Separation Using Visual Cues
Kang Zhang, Suyeon Lee, Arda Senocak +1
Cinematic Audio Source Separation (CASS) aims to decompose mixed film audio into speech, music, and sound effects, enabling applications like dubbing and remastering. Existing CASS…
DualTSR: Unified Dual-Diffusion Transformer for Scene Text Image Super-Resolution
Axi Niu, Kang Zhang, Qingsen Yan +3
Scene Text Image Super-Resolution (STISR) aims to restore high-resolution details in low-resolution text images, which is crucial for both human readability and machine recognition…
CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
Kangjie Zhang, Wenxuan Huang, Xin Zhou +9
Contrastive Language-Image Pre-training (CLIP) has achieved widely applications in various computer vision tasks, e.g., text-to-image generation, Image-Text retrieval and Image cap…
Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
Kang Zhang, Trung X. Pham, Suyeon Lee +3
We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike…
Physics Informed Distillation for Diffusion Models
Joshua Tian Jin Tee, Kang Zhang, Hee Suk Yoon +3
Diffusion models have recently emerged as a potent tool in generative modeling. However, their inherent iterative nature often results in sluggish image generation due to the requi…