collaborators

6 papers

cs.CV2025

SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery

Meng Cao, Xingyu Li, Xue Liu +2

Despite advancements in Multi-modal Large Language Models (MLLMs) for scene understanding, their performance on complex spatial reasoning tasks requiring mental simulation remains…

cs.CV2025

Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling

Meng Cao, Haokun Lin, Haoyuan Li +6

Spatial reasoning, the ability to understand and interpret the 3D structure of the world, is a critical yet underdeveloped capability in Multimodal Large Language Models (MLLMs). C…

cs.CV2025

Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning

Meng Cao, Haoze Zhao, Can Zhang +3

Large Vision-Language Models (LVLMs) have become powerful general-purpose assistants, yet their predictions often lack reliability and interpretability due to insufficient groundin…

cs.CV2025

Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

Meng Cao, Pengfei Hu, Yingyao Wang +12

Recent advancements in Large Video Language Models (LVLMs) have highlighted their potential for multi-modal understanding, yet evaluating their factual grounding in videos remains…

cs.CV2024

PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos

Meng Cao, Haoran Tang, Haoze Zhao +7

Recent advancements in video-based large language models (Video LLMs) have witnessed the emergence of diverse capabilities to reason and interpret dynamic visual content. Among the…

cs.CV2024

Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models

Meng Cao, Yuyang Liu, Yingfei Liu +6

Instruction tuning constitutes a prevalent technique for tailoring Large Vision Language Models (LVLMs) to meet individual task requirements. To date, most of the existing approach…