2 papers
cs.CV2026
Question-guided Visual Compression with Memory Feedback for Long-Term Video Understanding
Sosuke Yamao, Natsuki Miyahara, Yuankai Qi +1
In the context of long-term video understanding with large multimodal models, many frameworks have been proposed. Although transformer-based visual compressors and memory-augmented…
cs.CV2024
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs
Sosuke Yamao, Natsuki Miyahara, Yuki Harazono +1
With the increasing complexity of video data and the need for more efficient long-term temporal understanding, existing long-term video understanding methods often fail to accurate…