egocentric memory 1hierarchical memory 1multimodal streaming 1on-device inference 1personal AI assistants 1
From the 1 of 5 linked papers with an AI index.
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding
Jun Peng, Baiyang Song, Jie Li +4
Video understanding is often plagued by severe temporal redundancy, where processing dense frame sequences is both semantically inefficient and computationally expensive. This chal…
cs.CV2026
Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge
Yinsong Xu, Wei Jing, Liuxin Zhang +2
Understanding long-form egocentric videos remains challenging for multimodal large language models (MLLMs) due to limited context length and insufficient grounding of fine-grained…
cs.CV2026
Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
Tao Chen, Shaobo Ju, Qiong Wu +6
Due to excessive memory overhead, most Multimodal Large Language Models (MLLMs) can only process videos of limited frames. In this paper, we propose an effective and efficient para…