activity
20242026
most citedROOT: VLM based System for Indoor Scene Understanding and Beyond

1 citations · 1 across the 30 of their papers we have counts for

collaborators
Showing cs.CVShow all

23 papers · 1 filter

cs.CV2026

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering

Rongjian Gu, Wengang Zhou, Junyu Xiong +4

Multi-page document visual question answering requires locating sparse evidence at both the page and region levels. Existing approaches typically emphasize one level over the other…

cs.CV2026

Hierarchical Evidence-Driven Reasoning for Long Document Understanding

Junyu Xiong, Yonghui Wang, Rongjian Gu +4

Retrieval-Augmented Generation (RAG) streamlines long-document understanding by leveraging retrieval mechanisms to restrict input images to a highly curated subset. However, existi…

cs.CV2026

Revisiting Shadow Detection from a Vision-Language Perspective

Yonghui Wang, Shaokai Liu, Wengang Zhou +2

Shadow detection is commonly formulated as a vision-driven dense prediction problem, where models rely primarily on pixel-wise visual supervision to distinguish shadows from non-sh…

cs.CV2026

Geometry-Aware Dataset Condensation for Diffusion Model Training

Xiao Cui, Yulei Qin, Mo Zhu +3

Dataset condensation aims to construct compact datasets from real data via synthesis or selection. However, existing approaches are ill-suited for diffusion model training: synthet…

cs.CV2026

Edit-GRPO: A Locality-Preserving Policy Optimization Framework for Image Editing

Shaodong Xu, Zexian Li, Zhendong Wang +5

A fundamental challenge in image editing lies in preserving spatial locality: edits should improve targeted content without inadvertently altering surrounding regions. However, mos…

cs.CV2026

Beyond Point-Wise Matching: Structural Representation Alignment for Accelerating Diffusion Transformers

Shaodong Xu, Zhendong Wang, Litong Gong +4

Recent advances in Diffusion Transformers (DiTs) demonstrate that aligning noisy latent states with well-trained semantic features-as pioneered by Representation Alignment (REPA)-c…