Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Weighted Reverse Convolution for Feature Upsampling
Wentong Li, Zhiyuan Qi, Zichen Zhao +2
Pre-trained vision foundation models (VFMs) provide strong semantic representations, yet their patch-level features are inherently coarse, limiting their effectiveness on tasks req…
cs.CV2026
Prefix-Adaptive Block Diffusion for Efficient Document Recognition
Mingxu Chai, Ziyu Shen, Chenyu Liu +9
Block Diffusion Models (BDMs) support parallel generation, flexible-length output, and KV caching, making them promising for efficient document parsing. However, existing BDMs bind…