48 citations · 48 across the 7 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models
Mingxu Chai, Chenyu Liu, Ziyu Shen +7
Compared with typical vision-language tasks, document parsing places stronger demands on fine-grained visual perception. Existing vision-language model (VLM)-based parsing approach…
cs.CV2026
Prefix-Adaptive Block Diffusion for Efficient Document Recognition
Mingxu Chai, Ziyu Shen, Chenyu Liu +3
Block Diffusion Models (BDMs) support parallel generation, flexible-length output, and KV caching, making them promising for efficient document parsing. However, existing BDMs bind…
cs.CV2026
MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation
Changli Wu, Haodong Wang, Jiayi Ji +5
Most existing 3D referring expression segmentation (3DRES) methods rely on dense, high-quality point clouds, while real-world agents such as robots and mobile phones operate with o…