1 paper · 1 filter
Dongzhuoran Zhou, Yuqicheng Zhu, Yule Liu +5
Long-document visual question answering increasingly relies on VLM agents that retrieve candidate pages, inspect page images, write findings to working memory, and synthesize answe…