40 papers
HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering
Rongjian Gu, Wengang Zhou, Junyu Xiong +4
Multi-page document visual question answering requires locating sparse evidence at both the page and region levels. Existing approaches typically emphasize one level over the other…
Hierarchical Evidence-Driven Reasoning for Long Document Understanding
Junyu Xiong, Yonghui Wang, Rongjian Gu +4
Retrieval-Augmented Generation (RAG) streamlines long-document understanding by leveraging retrieval mechanisms to restrict input images to a highly curated subset. However, existi…
Revisiting Shadow Detection from a Vision-Language Perspective
Yonghui Wang, Shaokai Liu, Wengang Zhou +2
Shadow detection is commonly formulated as a vision-driven dense prediction problem, where models rely primarily on pixel-wise visual supervision to distinguish shadows from non-sh…
Geometry-Aware Dataset Condensation for Diffusion Model Training
Xiao Cui, Yulei Qin, Mo Zhu +3
Dataset condensation aims to construct compact datasets from real data via synthesis or selection. However, existing approaches are ill-suited for diffusion model training: synthet…
Self-supervised Hierarchical Visual Reasoning with World Model
Yuanfei Xu, Lin Liu, Wengang Zhou +2
3D open-world environments with adversarial opponents remain a core challenge for reinforcement learning due to their vast state spaces. Effective reasoning representations are ess…
Edit-GRPO: A Locality-Preserving Policy Optimization Framework for Image Editing
Shaodong Xu, Zexian Li, Zhendong Wang +5
A fundamental challenge in image editing lies in preserving spatial locality: edits should improve targeted content without inadvertently altering surrounding regions. However, mos…