11 papers
Are All Tokens Necessary for Visual Place Recognition? An Empirical Study of Token Reduction for Efficient Inference
Tong Jin, Yunpeng Liu, Shuyu Hu +4
Recent visual place recognition (VPR) methods based on vision transformers, particularly foundation models, have achieved remarkable recognition performance. However, these models…
Selectivity Drives Efficiency: Dataset Pruning for Visual Place Recognition
Tong Jin, Yunpeng Liu, Shuyu Hu +3
The paper introduces a place-wise dataset pruning method for visual place recognition that selects informative locations using intra-place diversity and inter-place similarity metr…
Online Reasoning Video Object Segmentation
Jinyuan Liu, Yang Wang, Zeyu Zhao +3
Reasoning video object segmentation predicts pixel-level masks in videos from natural-language queries that may involve implicit and temporally grounded references. However, existi…
Data-Efficient Surgical Phase Segmentation in Small-Incision Cataract Surgery: A Controlled Study of Vision Foundation Models
Lincoln Spencer, Song Wang, Chen Chen
Surgical phase segmentation is central to computer-assisted surgery, yet robust models remain difficult to develop when labeled surgical videos are scarce. We study data-efficient…
AdaSFormer: Adaptive Serialized Transformers for Monocular Semantic Scene Completion from Indoor Environments
Xuzhi Wang, Xinran Wu, Song Wang +2
Indoor monocular semantic scene completion (MSSC) is notably more challenging than its outdoor counterpart due to complex spatial layouts and severe occlusions. While transformers…
ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps
Sicheng Feng, Song Wang, Shuyi Ouyang +5
Multimodal large language models (MLLMs) have demonstrated significant progress in semantic scene understanding and text-image alignment, with reasoning variants enhancing performa…