Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models
Mingxu Chai, Chenyu Liu, Ziyu Shen +7
Compared with typical vision-language tasks, document parsing places stronger demands on fine-grained visual perception. Existing vision-language model (VLM)-based parsing approach…
cs.CV2026
Prefix-Adaptive Block Diffusion for Efficient Document Recognition
Mingxu Chai, Ziyu Shen, Chenyu Liu +3
Block Diffusion Models (BDMs) support parallel generation, flexible-length output, and KV caching, making them promising for efficient document parsing. However, existing BDMs bind…
cs.CV2025
ACD-CLIP: Decoupling Representation and Dynamic Fusion for Zero-Shot Anomaly Detection
Ke Ma, Jun Long, Hongxiao Fei +3
Pre-trained Vision-Language Models (VLMs) struggle with Zero-Shot Anomaly Detection (ZSAD) due to a critical adaptation gap: they lack the local inductive biases required for dense…