4 papers
Revisiting Shadow Detection from a Vision-Language Perspective
Yonghui Wang, Shaokai Liu, Wengang Zhou +2
Shadow detection is commonly formulated as a vision-driven dense prediction problem, where models rely primarily on pixel-wise visual supervision to distinguish shadows from non-sh…
Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters
Gengluo Li, Shangpin Peng, Xingyu Wan +16
Vision Large Language Models (VLLMs) have achieved remarkable success in modern text-rich visual understanding. However, their perceptual robustness in the face of the continuous m…
BookNet: Book Image Rectification via Cross-Page Attention Network
Shaokai Liu, Hao Feng, Bozhi Luan +3
Book image rectification presents unique challenges in document image processing due to complex geometric distortions from binding constraints, where left and right pages exhibit d…
Multi-Cue Adaptive Visual Token Pruning for Large Vision-Language Models
Bozhi Luan, Wengang Zhou, Hao Feng +3
As the computational needs of Large Vision-Language Models (LVLMs) increase, visual token pruning has proven effective in improving inference speed and memory efficiency. Tradition…