most citedStructured Attention Matters to Multimodal LLMs in Document Understanding

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CV2025

SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models

Hongxing Li, Dingming Li, Zixuan Wang +7

Spatial reasoning remains a fundamental challenge for Vision-Language Models (VLMs), with current approaches struggling to achieve robust performance despite recent advances. We id…

cs.CV2025

FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning

Haonan Ge, Yiwei Wang, Kai-Wei Chang +2

Current video understanding models rely on fixed frame sampling strategies, processing predetermined visual inputs regardless of the specific reasoning requirements of each questio…

cs.AI2025

RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation

Hang Wu, Yujun Cai, Haonan Ge +3

Cinematography understanding refers to the ability to recognize not only the visual content of a scene but also the cinematic techniques that shape narrative meaning. This capabili…

cs.AI2025

DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning

Hang Wu, Hongkai Chen, Yujun Cai +4

Grounding natural language queries in graphical user interfaces (GUIs) poses unique challenges due to the diversity of visual elements, spatial clutter, and the ambiguity of langua…

cs.CL20251 cited

Structured Attention Matters to Multimodal LLMs in Document Understanding

Chang Liu, Hongkai Chen, Yujun Cai +4

Document understanding remains a significant challenge for multimodal large language models (MLLMs). While previous research has primarily focused on locating evidence pages throug…

cs.CV2025

STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation

Jiamin Wang, Yichen Yao, Xiang Feng +5

The generation of temporally consistent, high-fidelity driving videos over extended horizons presents a fundamental challenge in autonomous driving world modeling. Existing approac…