activity
20242026
collaborators

6 papers

cs.CV2026

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering

Rongjian Gu, Wengang Zhou, Junyu Xiong +4

Multi-page document visual question answering requires locating sparse evidence at both the page and region levels. Existing approaches typically emphasize one level over the other…

cs.CV2026

Hierarchical Evidence-Driven Reasoning for Long Document Understanding

Junyu Xiong, Yonghui Wang, Rongjian Gu +4

Retrieval-Augmented Generation (RAG) streamlines long-document understanding by leveraging retrieval mechanisms to restrict input images to a highly curated subset. However, existi…

cs.CV2026

Revisiting Shadow Detection from a Vision-Language Perspective

Yonghui Wang, Shaokai Liu, Wengang Zhou +2

Shadow detection is commonly formulated as a vision-driven dense prediction problem, where models rely primarily on pixel-wise visual supervision to distinguish shadows from non-sh…

cs.CV2025

DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding

Junyu Xiong, Yonghui Wang, Weichao Zhao +4

Understanding multi-page documents poses a significant challenge for multimodal large language models (MLLMs), as it requires fine-grained visual comprehension and multi-hop reason…

cs.CV2025

Text-to-3D Generation by 2D Editing

Haoran Li, Yuli Tian, Yonghui Wang +4

Distilling 3D representations from pretrained 2D diffusion models is essential for 3D creative applications across gaming, film, and interior design. Current SDS-based methods are…

cs.CV2024

ROOT: VLM based System for Indoor Scene Understanding and Beyond

Yonghui Wang, Shi-Yong Chen, Zhenxing Zhou +4

Recently, Vision Language Models (VLMs) have experienced significant advancements, yet these models still face challenges in spatial hierarchical reasoning within indoor scenes. In…