5 papers
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
Lingxiao Li, Yifan Wang, Xinyan Gao +3
Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential in multimodal large language mo…
Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning
Xinyan Gao, Haoran Hao, Xiangyu Yue
The rapid development of pretrained foundation models has enabled more general image segmentation. Multimodal large language models (MLLMs) have been widely explored for image segm…
-WM: A Unified Video-Action World Model for Robotic Manipulation
Pengfei Zhou, Shengcong Chen, Di Chen +17
Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execution. We present -World…
A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning
Xiaoda Yang, Shuai Yang, Can Wang +9
Vision-Language Models (VLMs) have made significant strides in static image understanding but continue to face critical hurdles in spatiotemporal reasoning. A major bottleneck is "…
3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
Xiaoye Wang, Chen Tang, Xiangyu Yue +1
This paper addresses the challenge of training a single network to jointly perform multiple dense prediction tasks, such as segmentation and depth estimation, i.e., multi-task lear…