collaborators

51 papers

cs.CV2026

Thinking with Anchors: Grounded and Efficient Document Reasoning

Sichen Zhu, Yuchen Zhu, Wenzhuo Xu +13

Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region seman…

cs.CL2026

Lost in a Single Vector: Improving Long-Document Retrieval with Chunk Evidence Aggregation

Shanshan Lyu, Yiwei Wang, Yujun Cai +2

Dense retrieval ranks one query vector against one document vector. On long documents, this interface can fail when a short but decisive span is weakened during document encoding b…

cs.CL2026

Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language Models

Boyan Han, Yiwei Wang, Yi Song +2

Diffusion large language models (dLLMs) offer bidirectional attention and parallel generation, enabling them to exploit global context and naturally support format-constrained task…

cs.CV2026

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding

Hang Wu, Sherin Mary Mathews, Yujun Cai +2

Online streaming video understanding requires models to process continuous visual inputs and respond to user queries in real time, where the unbounded stream and unpredictable quer…

cs.CV2026

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG

Honghao Fu, Miao Xu, Yiwei Wang +3

Scaling multimodal large language models (MLLMs) to long videos is constrained by limited context windows. While retrieval-augmented generation (RAG) is a promising remedy by organ…

cs.CV2026

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning

Hang Wu, Yujun Cai, Zehao Li +4

Understanding camera dynamics is a fundamental pillar of video spatial intelligence. However, existing multimodal models predominantly treat this task as a black-box classification…