9 papers
COMBINER: Composed Image Retrieval Guided by Attribute-based Neighbor Relations
Zixu Li, Yupeng Hu, Zhiwei Chen +3
Composed Image Retrieval (CIR) represents a challenging retrieval task that targets locating specific images through multimodal inputs. Despite recent progress in CIR techniques, p…
STABLE: Efficient Hybrid Nearest Neighbor Search via Magnitude-Uniformity and Cardinality-Robustness
Qianyun Yang, Zhiwei Chen, Yupeng Hu +3
Hybrid Approximate Nearest Neighbor Search (Hybrid ANNS) is a foundational search technology for large-scale heterogeneous data and has gained significant attention in both academi…
ELIP: Efficient Discriminative Language-Image Pre-training with Fewer Vision Tokens
Yangyang Guo, Haoyu Zhang, Yongkang Wong +2
Learning a versatile language-image model is computationally prohibitive under a limited computing budget. This paper delves into the \emph{efficient language-image pre-training},…
Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
Qiaohui Chu, Haoyu Zhang, Meng Liu +3
Long-term action anticipation from egocentric video is critical for applications such as human-computer interaction and assistive technologies, where anticipating user intent enabl…
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog
Haoyu Zhang, Meng Liu, Yisen Feng +3
In contrast to conventional visual question answering, video-grounded dialog necessitates a profound understanding of both dialog history and video content for accurate response ge…
MaeFuse: Transferring Omni Features with Pretrained Masked Autoencoders for Infrared and Visible Image Fusion via Guided Training
Jiayang Li, Junjun Jiang, Pengwei Liang +2
In this paper, we introduce MaeFuse, a novel autoencoder model designed for Infrared and Visible Image Fusion (IVIF). The existing approaches for image fusion often rely on trainin…