Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Visual Information Extraction from Documents via Classification-Guided Large Vision-Language Models
Huafu Li, Guo Chen, Jia Xia +5
Visual information extraction (VIE) from visually rich documents remains challenging due to high layout variability and real-world impairments. Existing methods typically rely on s…
cs.CV2024
Retrieval-Augmented Egocentric Video Captioning
Jilan Xu, Yifei Huang, Junlin Hou +4
Understanding human actions from videos of first-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only, while ov…
cs.CV2024
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Kunchang Li, Yali Wang, Yinan He +9
With the rapid development of Multi-modal Large Language Models (MLLMs), a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these…