13 papers
SC-Diff: Semantically Calibrated Diffusion for Visible-to-Infrared Image Translation
Junyin Zhang, Siyu Huang, Jianxiong Ye +4
Visible-to-infrared image translation provides a practical way to expand infrared training data using abundant visible images. Diffusion models are promising for this task because…
Registration-Grounded Spectral Fusion for Unregistered WLI/NBI Endoscopic Lesion Segmentation
Pengyu Jie, Wanquan Liu, Rui He +5
The paper proposes a reliability‑aware framework that first aligns features from white‑light and narrow‑band endoscopic images and then fuses them in a complex‑valued representatio…
Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding
Biao Tang, Xu Chen, Shuxiang Gou +3
Long-video understanding remains challenging for multimodal large language models, because temporally extended videos often contain thousands of frames and are therefore expensive…
IV-tuning: Parameter-Efficient Transfer Learning for Infrared-Visible Tasks
Yaming Zhang, Chenqiang Gao, Fangcen Liu +4
Existing infrared and visible (IR-VIS) methods inherit the general representations of Pre-trained Visual Models (PVMs) to facilitate complementary learning. However, our analysis i…
WildGHand: Learning Anti-Perturbation Gaussian Hand Avatars from Monocular In-the-Wild Videos
Hanhui Li, Xuan Huang, Wanquan Liu +5
Despite recent progress in 3D hand reconstruction from monocular videos, most existing methods rely on data captured in well-controlled environments and therefore degrade in real-w…
Are Dense Labels Always Necessary for 3D Object Detection from Point Cloud?
Chenqiang Gao, Chuandong Liu, Jun Shu +5
Current state-of-the-art (SOTA) 3D object detection methods often require a large amount of 3D bounding box annotations for training. However, collecting such large-scale densely-s…