papers

Publications (10)

cs.CV2026

VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion

Linfeng Tang, Yeda Wang, Meiqi Gong +7

Compared to images, videos better reflect real-world acquisition and possess valuable temporal cues. However, existing multi-sensor fusion research predominantly integrates complem…

cs.CV2025

Deep Learning Reforms Image Matching: A Survey and Outlook

Shihua Zhang, Zizhuo Li, Kaining Zhang +5

Image matching, which establishes correspondences between two-view images to recover 3D structure and camera geometry, serves as a cornerstone in computer vision and underpins a wi…

cs.CV2026

AnyMatch: Supercharging Universal Multi-Modal Image Matching with Large-Scale Single-View Images

Meng Yang, Zizhuo Li, Linfeng Tang +2

Multi-modal image matching is essential for visual localization and multi-sensor fusion, but it is hindered by the scarcity of large-scale training data with precise geometric anno…

cs.CV2025

CoMatch: Dynamic Covisibility-Aware Transformer for Bilateral Subpixel-Level Semi-Dense Image Matching

Zizhuo Li, Yifan Lu, Linfeng Tang +2

This prospective study proposes CoMatch, a novel semi-dense image matcher with dynamic covisibility awareness and bilateral subpixel accuracy. Firstly, observing that modeling cont…

cs.CV2024

Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion

Xunpeng Yi, Han Xu, Hao Zhang +2

Image fusion aims to combine information from different source images to create a comprehensively representative image. Existing fusion methods are typically helpless in dealing wi…

cs.CV2025

TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration

Meiqi Gong, Hao Zhang, Xunpeng Yi +2

Existing multi-modal fusion methods typically apply static frame-based image fusion techniques directly to video fusion tasks, neglecting inherent temporal dependencies and leading…