4 papers
Perceptio: Perception Enhanced Vision Language Models via Spatial Token Generation
Yuchen Li, Amanmeet Garg, Shalini Chaudhuri +2
Large Vision Language Models (LVLMs) excel at semantic understanding but struggle with fine grained spatial grounding, as the model must implicitly infer complex geometry without e…
What Happens Next? Next Scene Prediction with a Unified Video Model
Xinjie Li, Zhimin Chen, Rui Zhao +3
Recent unified models for joint understanding and generation have significantly advanced visual generation capabilities. However, their focus on conventional tasks like text-to-vid…
VIDEOP2R: Video Understanding from Perception to Reasoning
Yifan Jiang, Yueying Wang, Rui Zhao +4
Reinforcement fine-tuning (RFT), a two-stage framework consisting of supervised fine-tuning (SFT) and reinforcement learning (RL) has shown promising results on improving reasoning…
Causal Reasoning Elicits Controllable 3D Scene Generation
Shen Chen, Ruiyu Zhao, Jiale Zhou +3
Existing 3D scene generation methods often struggle to model the complex logical dependencies and physical constraints between objects, limiting their ability to adapt to dynamic a…