6 papers
HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering
Rongjian Gu, Wengang Zhou, Junyu Xiong +4
Multi-page document visual question answering requires locating sparse evidence at both the page and region levels. Existing approaches typically emphasize one level over the other…
Hierarchical Evidence-Driven Reasoning for Long Document Understanding
Junyu Xiong, Yonghui Wang, Rongjian Gu +4
Retrieval-Augmented Generation (RAG) streamlines long-document understanding by leveraging retrieval mechanisms to restrict input images to a highly curated subset. However, existi…
Revisiting Shadow Detection from a Vision-Language Perspective
Yonghui Wang, Shaokai Liu, Wengang Zhou +2
Shadow detection is commonly formulated as a vision-driven dense prediction problem, where models rely primarily on pixel-wise visual supervision to distinguish shadows from non-sh…
DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding
Junyu Xiong, Yonghui Wang, Weichao Zhao +4
Understanding multi-page documents poses a significant challenge for multimodal large language models (MLLMs), as it requires fine-grained visual comprehension and multi-hop reason…
Text-to-3D Generation by 2D Editing
Haoran Li, Yuli Tian, Yonghui Wang +4
Distilling 3D representations from pretrained 2D diffusion models is essential for 3D creative applications across gaming, film, and interior design. Current SDS-based methods are…
ROOT: VLM based System for Indoor Scene Understanding and Beyond
Yonghui Wang, Shi-Yong Chen, Zhenxing Zhou +4
Recently, Vision Language Models (VLMs) have experienced significant advancements, yet these models still face challenges in spatial hierarchical reasoning within indoor scenes. In…