1 citations · 1 across the 4 of their papers we have counts for
6 papers
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
Hongxing Li, Dingming Li, Zixuan Wang +7
Spatial reasoning remains a fundamental challenge for Vision-Language Models (VLMs), with current approaches struggling to achieve robust performance despite recent advances. We id…
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
Haonan Ge, Yiwei Wang, Kai-Wei Chang +2
Current video understanding models rely on fixed frame sampling strategies, processing predetermined visual inputs regardless of the specific reasoning requirements of each questio…
RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
Hang Wu, Yujun Cai, Haonan Ge +3
Cinematography understanding refers to the ability to recognize not only the visual content of a scene but also the cinematic techniques that shape narrative meaning. This capabili…
DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
Hang Wu, Hongkai Chen, Yujun Cai +4
Grounding natural language queries in graphical user interfaces (GUIs) poses unique challenges due to the diversity of visual elements, spatial clutter, and the ambiguity of langua…
Structured Attention Matters to Multimodal LLMs in Document Understanding
Chang Liu, Hongkai Chen, Yujun Cai +4
Document understanding remains a significant challenge for multimodal large language models (MLLMs). While previous research has primarily focused on locating evidence pages throug…
STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation
Jiamin Wang, Yichen Yao, Xiang Feng +5
The generation of temporally consistent, high-fidelity driving videos over extended horizons presents a fundamental challenge in autonomous driving world modeling. Existing approac…