collaborators

5 papers

cs.CV2026

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning

Lingxiao Li, Yifan Wang, Xinyan Gao +3

Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential in multimodal large language mo…

cs.CV2026

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning

Xinyan Gao, Haoran Hao, Xiangyu Yue

The rapid development of pretrained foundation models has enabled more general image segmentation. Multimodal large language models (MLLMs) have been widely explored for image segm…

cs.RO2026

-WM: A Unified Video-Action World Model for Robotic Manipulation

Pengfei Zhou, Shengcong Chen, Di Chen +17

Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execution. We present -World…

cs.AI2026

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning

Xiaoda Yang, Shuai Yang, Can Wang +9

Vision-Language Models (VLMs) have made significant strides in static image understanding but continue to face critical hurdles in spatiotemporal reasoning. A major bottleneck is "…

cs.CV2025

3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding

Xiaoye Wang, Chen Tang, Xiangyu Yue +1

This paper addresses the challenge of training a single network to jointly perform multiple dense prediction tasks, such as segmentation and depth estimation, i.e., multi-task lear…