Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
ViCrop-Det: Spatial Attention Entropy Guided Cropping for Training-Free Small-Object Detection
Hui Wang, Hongze Li, Wei Chen +1
Transformer-based architectures have established a dominant paradigm in global semantic perception; however, they remain fundamentally constrained by the profound spatial heterogen…
cs.CV2024
Do Current Video LLMs Have Strong OCR Abilities? A Preliminary Study
Yulin Fei, Yuhui Gao, Xingyuan Xian +3
With the rise of multimodal large language models, accurately extracting and understanding textual information from video content, referred to as video based optical character reco…