4 papers
GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning
Zhaochen Liu, Limeng Qiao, Guanglu Wan +1
Multimodal large language models (MLLMs) have exhibited remarkable performance in various visual tasks, yet still struggle with spatial reasoning. Recent efforts mitigate this by i…
Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
Zhaochen Liu, Kaiwen Gao, Shuyi Liang +4
Occlusion perception, a critical foundation for human-level spatial understanding, embodies the challenge of integrating visual recognition and reasoning. Though multimodal large l…
Amodal Segmentation for Laparoscopic Surgery Video Instruments
Ruohua Shi, Zhaochen Liu, Lingyu Duan +1
Segmentation of surgical instruments is crucial for enhancing surgeon performance and ensuring patient safety. Conventional techniques such as binary, semantic, and instance segmen…
PLUG: Revisiting Amodal Segmentation with Foundation Model and Hierarchical Focus
Zhaochen Liu, Limeng Qiao, Xiangxiang Chu +1
Aiming to predict the complete shapes of partially occluded objects, amodal segmentation is an important step towards visual intelligence. With crucial significance, practical prio…