116 citations · 125 across the 6 of their papers we have counts for
7 papers · 1 filter
GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience
Sitong Gong, Caixin Kang, Tianyu Yan +7
A wearable assistant should both answer questions about its visual history and recognize when that history is useful to the present situation. Existing video-memory systems primari…
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
Gong Sitong, Tianyu Yan, Caixin Kang +6
When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proa…
Deep Boosting Learning: A Brand-new Cooperative Approach for Image-Text Matching
Haiwen Diao, Ying Zhang, Shang Gao +2
Image-text matching remains a challenging task due to heterogeneous semantic diversity across modalities and insufficient distance separability within triplets. Different from prev…
Plug-and-Play Regulators for Image-Text Matching
Haiwen Diao, Ying Zhang, Wei Liu +2
Exploiting fine-grained correspondence and visual-semantic alignments has shown great potential in image-text matching. Generally, recent approaches first employ a cross-modal atte…
OMG: Observe Multiple Granularities for Natural Language-Based Vehicle Retrieval
Yunhao Du, Binyu Zhang, Xiangning Ruan +3
Retrieving tracked-vehicles by natural language descriptions plays a critical role in smart city construction. It aims to find the best match for the given texts from a set of trac…
Visible-Thermal UAV Tracking: A Large-Scale Benchmark and New Baseline
Pengyu Zhang, Jie Zhao, Dong Wang +2
With the popularity of multi-modal sensors, visible-thermal (RGB-T) object tracking is to achieve robust performance and wider application scenarios with the guidance of objects' t…