4 papers · 1 filter
HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering
Rongjian Gu, Wengang Zhou, Junyu Xiong +4
Multi-page document visual question answering requires locating sparse evidence at both the page and region levels. Existing approaches typically emphasize one level over the other…
Hierarchical Evidence-Driven Reasoning for Long Document Understanding
Junyu Xiong, Yonghui Wang, Rongjian Gu +4
Retrieval-Augmented Generation (RAG) streamlines long-document understanding by leveraging retrieval mechanisms to restrict input images to a highly curated subset. However, existi…
DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding
Junyu Xiong, Yonghui Wang, Weichao Zhao +4
Understanding multi-page documents poses a significant challenge for multimodal large language models (MLLMs), as it requires fine-grained visual comprehension and multi-hop reason…
M-LLM Based Video Frame Selection for Efficient Video Understanding
Kai Hu, Feng Gao, Xiaohan Nie +8
Recent advances in Multi-Modal Large Language Models (M-LLMs) show promising results in video reasoning. Popular Multi-Modal Large Language Model (M-LLM) frameworks usually apply n…