Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
Junbin Xiao, Jiajun Chen, Tianxiang Sun +2
Long streaming video QA remains challenging due to growing visual tokens and limited reasoning length of large language models (LLMs). KV-caching stores the Key-Value (KV) of the h…
cs.CV2026
SAVER: Selective As-Needed Vision Evidence for Multimodal Information Extraction
Miaobo Hu, Shuhao Hu, Bokun Wang +5
Multimodal IE in social media is difficult because a post may attach multiple images that are weakly related, redundant, or even misleading with respect to the text. In this settin…
cs.CV2026
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
Wei Wang, Yuqian Yuan, Tianwei Lin +4
Spatial intelligence requires multimodal large language models (MLLMs) to move beyond single-view perception and reason consistently about objects, visibility, geometry, and intera…