6 papers
Do MLLMs Exhibit Human-like Perceptual Behaviors? HVSBench: A Benchmark for MLLM Alignment with Human Perceptual Behavior
Jiaying Lin, Shuquan Ye, Dan Xu +2
While Multimodal Large Language Models (MLLMs) excel at many vision tasks, it is unknown if they exhibit human-like perceptual behaviors. To evaluate this, we introduce HVSBench, t…
OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene Understanding
Youjun Zhao, Jiaying Lin, Shuquan Ye +2
Open-vocabulary 3D scene understanding (OV-3D) aims to localize and classify novel objects beyond the closed set of object classes. However, existing approaches and benchmarks prim…
MirrorMamba: Towards Scalable and Robust Mirror Detection in Videos
Rui Song, Jiaying Lin, Rynson W. H. Lau
Video mirror detection has received significant research attention, yet existing methods suffer from limited performance and robustness. These approaches often over-rely on single,…
Hierarchical Cross-Modal Alignment for Open-Vocabulary 3D Object Detection
Youjun Zhao, Jiaying Lin, Rynson W. H. Lau
Open-vocabulary 3D object detection (OV-3DOD) aims at localizing and classifying novel objects beyond closed sets. The recent success of vision-language models (VLMs) has demonstra…
Leveraging RGB-D Data with Cross-Modal Context Mining for Glass Surface Detection
Jiaying Lin, Yuen-Hei Yeung, Shuquan Ye +1
Glass surfaces are becoming increasingly ubiquitous as modern buildings tend to use a lot of glass panels. This, however, poses substantial challenges to the operations of autonomo…
Boosting Weakly-Supervised Referring Image Segmentation via Progressive Comprehension
Zaiquan Yang, Yuhao Liu, Jiaying Lin +2
This paper explores the weakly-supervised referring image segmentation (WRIS) problem, and focuses on a challenging setup where target localization is learned directly from image-t…