9 papers
Unleashing the Power of Text: Text-Guided Flow Matching for Image Fusion under Complex Degradations
Axi Niu, Jieheng Li, Kang Zhang +3
Infrared-visible image fusion under realistic degradation scenarios is a challenging task, as degradations not only cause a loss of reliable modality-specific information in observ…
Cinematic Audio Source Separation Using Visual Cues
Kang Zhang, Suyeon Lee, Arda Senocak +1
Cinematic Audio Source Separation (CASS) aims to decompose mixed film audio into speech, music, and sound effects, enabling applications like dubbing and remastering. Existing CASS…
DualTSR: Unified Dual-Diffusion Transformer for Scene Text Image Super-Resolution
Axi Niu, Kang Zhang, Qingsen Yan +3
Scene Text Image Super-Resolution (STISR) aims to restore high-resolution details in low-resolution text images, which is crucial for both human readability and machine recognition…
CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
Kangjie Zhang, Wenxuan Huang, Xin Zhou +9
Contrastive Language-Image Pre-training (CLIP) has achieved widely applications in various computer vision tasks, e.g., text-to-image generation, Image-Text retrieval and Image cap…
Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation
Kang Zhang, Trung X. Pham, Suyeon Lee +3
We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike…
E-MD3C: Taming Masked Diffusion Transformers for Efficient Zero-Shot Object Customization
Trung X. Pham, Zhang Kang, Ji Woo Hong +2
We propose E-MD3C (fficient asked iffusion Transformer with Disentangled onditions and ompact $\underline…