51 papers
Thinking with Anchors: Grounded and Efficient Document Reasoning
Sichen Zhu, Yuchen Zhu, Wenzhuo Xu +13
Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region seman…
Lost in a Single Vector: Improving Long-Document Retrieval with Chunk Evidence Aggregation
Shanshan Lyu, Yiwei Wang, Yujun Cai +2
Dense retrieval ranks one query vector against one document vector. On long documents, this interface can fail when a short but decisive span is weakened during document encoding b…
Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language Models
Boyan Han, Yiwei Wang, Yi Song +2
Diffusion large language models (dLLMs) offer bidirectional attention and parallel generation, enabling them to exploit global context and naturally support format-constrained task…
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
Hang Wu, Sherin Mary Mathews, Yujun Cai +2
Online streaming video understanding requires models to process continuous visual inputs and respond to user queries in real time, where the unbounded stream and unpredictable quer…
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
Honghao Fu, Miao Xu, Yiwei Wang +3
Scaling multimodal large language models (MLLMs) to long videos is constrained by limited context windows. While retrieval-augmented generation (RAG) is a promising remedy by organ…
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
Hang Wu, Yujun Cai, Zehao Li +4
Understanding camera dynamics is a fundamental pillar of video spatial intelligence. However, existing multimodal models predominantly treat this task as a black-box classification…