8 papers
AnyMatch: Supercharging Universal Multi-Modal Image Matching with Large-Scale Single-View Images
Meng Yang, Zizhuo Li, Linfeng Tang +2
Multi-modal image matching is essential for visual localization and multi-sensor fusion, but it is hindered by the scarcity of large-scale training data with precise geometric anno…
Efficient Sparse-to-Dense Visual Localization via Compact Gaussian Scene Representation and Accelerated Dense Pose Estimation
Zizhuo Li, Songchu Deng, Linfeng Tang +1
This letter presents LiteLoc, a novel and efficient localizer built on 3D Gaussian Splatting (3DGS). The previous state-of-the-art (SoTA) sparse-to-dense localizer, STDLoc, has sho…
VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion
Linfeng Tang, Yeda Wang, Meiqi Gong +7
Compared to images, videos better reflect real-world acquisition and possess valuable temporal cues. However, existing multi-sensor fusion research predominantly integrates complem…
ControlFusion: A Controllable Image Fusion Framework with Language-Vision Degradation Prompts
Linfeng Tang, Yeda Wang, Zhanchuan Cai +2
Current image fusion methods struggle to address the composite degradations encountered in real-world imaging scenarios and lack the flexibility to accommodate user-specific requir…
TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration
Meiqi Gong, Hao Zhang, Xunpeng Yi +2
Existing multi-modal fusion methods typically apply static frame-based image fusion techniques directly to video fusion tasks, neglecting inherent temporal dependencies and leading…
Deep Learning Reforms Image Matching: A Survey and Outlook
Shihua Zhang, Zizhuo Li, Kaining Zhang +5
Image matching, which establishes correspondences between two-view images to recover 3D structure and camera geometry, serves as a cornerstone in computer vision and underpins a wi…