6 citations · 12 across the 47 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering
Qing Yang, Pengcheng Huang, Xinze Li +6
Long-video question answering remains challenging for Vision-Language Models (VLMs), as answer-relevant evidence is often sparse, transient, and temporally dispersed across lengthy…
cs.CV2026
Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation
Yichen Zhang, Da Peng, Zonghao Guo +19
A recent cutting-edge topic in multimodal modeling is to unify visual comprehension and generation within a single model. However, the two tasks demand mismatched decoding regimes…