4 papers
OrientSAM: Mitigating Camera-Centric Shortcut in Multimodal Spatial Reasoning via Orientation-Aware Spatial Alignment
Wenxiao Fan, Hang Yin, Kan Li
Multimodal large language models (MLLMs) still struggle with spatial reasoning that requires perspective transformation. In particular, they often rely on camera-centric cues rathe…
Stitch and Tell: A Structured Multimodal Data Augmentation Method for Spatial Understanding
Hang Yin, Xiaomin He, PeiWen Yuan +5
Existing vision-language models often suffer from spatial hallucinations, i.e., generating incorrect descriptions about the relative positions of objects in an image. We argue that…
Generative Video Semantic Communication via Multimodal Semantic Fusion with Large Model
Hang Yin, Li Qiao, Yu Ma +4
Despite significant advancements in traditional syntactic communications based on Shannon's theory, these methods struggle to meet the requirements of 6G immersive communications,…
Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning
Hang Yin, Zhifeng Lin, Xin Liu +2
Direction reasoning is essential for intelligent systems to understand the real world. While existing work focuses primarily on spatial reasoning, compass direction reasoning remai…