Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
SAVER: Selective As-Needed Vision Evidence for Multimodal Information Extraction
Miaobo Hu, Shuhao Hu, Bokun Wang +5
Multimodal IE in social media is difficult because a post may attach multiple images that are weakly related, redundant, or even misleading with respect to the text. In this settin…
cs.CV2026
Video Compression Meets Video Generation: Latent Inter-Frame Pruning with Attention Recovery
Dennis Menn, Yuedong Yang, Bokun Wang +6
Current video generation models suffer from high computational latency, making real-time applications prohibitively costly. In this paper, we address this limitation by exploiting…