Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams
Xiaokang Ma, Yifan Sun, Zhihong Jin +6
Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agen…
cs.CV2025
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective
Lei Lei, Jie Gu, Xiaokang Ma +3
Existing Multimodal Large Language Models (MLLMs) process a large number of visual tokens, leading to significant computational costs and inefficiency. Instruction-related visual t…