From the 1 of 4 linked papers with an AI index.
4 papers
ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams
Xiaokang Ma, Yifan Sun, Zhihong Jin +6
The paper introduces ReflectWorld-MM, an entity-oriented multimodal memory architecture that processes continuous video streams, stores observations in a hierarchical long‑term mem…
Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective
Lei Lei, Jie Gu, Xiaokang Ma +3
Existing Multimodal Large Language Models (MLLMs) process a large number of visual tokens, leading to significant computational costs and inefficiency. Instruction-related visual t…
Egocentric Instruction-oriented Affordance Prediction via Large Multimodal Model
Bokai Ji, Jie Gu, Xiaokang Ma +3
Affordance is crucial for intelligent robots in the context of object manipulation. In this paper, we argue that affordance should be task-/instruction-dependent, which is overlook…
Stimulating Imagination: Towards General-purpose "Something Something Placement"
Jianyang Wu, Jie Gu, Xiaokang Ma +3
General-purpose object placement is a fundamental capability of an intelligent generalist robot: being capable of rearranging objects following precise human instructions even in n…