works on

From the 2 of 7 linked papers with an AI index.

collaborators

7 papers

cs.RO2026

Move-Then-Operate: Behavioral Phasing for Human-Like Robotic Manipulation

Haoming Xu, Lei Lei, Jie Gu +3

The paper introduces Move-Then-Operate, a vision‑language framework that splits robotic manipulation into a coarse relocation phase and a contact‑critical interaction phase using a…

cs.CV2026

ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams

Xiaokang Ma, Yifan Sun, Zhihong Jin +6

The paper introduces ReflectWorld-MM, an entity-oriented multimodal memory architecture that processes continuous video streams, stores observations in a hierarchical long‑term mem…

cs.CV2026

DeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from Videos

Can Li, Zhoujian Li, Ren Li +4

World models for deformable objects should recover not only geometry and appearance, but also underlying physical dynamics, interaction grounding, and material behavior. Learning s…

cs.CV2026

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective

Lei Lei, Jie Gu, Xiaokang Ma +3

Existing Multimodal Large Language Models (MLLMs) process a large number of visual tokens, leading to significant computational costs and inefficiency. Instruction-related visual t…

cs.CV2026

Gaussian Sequences with Multi-Scale Dynamics for 4D Reconstruction from Monocular Casual Videos

Can Li, Jie Gu, Jingmin Chen +2

Understanding dynamic scenes from casual videos is critical for scalable robot learning, yet four-dimensional (4D) reconstruction under strictly monocular settings remains highly i…

cs.RO2025

Egocentric Instruction-oriented Affordance Prediction via Large Multimodal Model

Bokai Ji, Jie Gu, Xiaokang Ma +3

Affordance is crucial for intelligent robots in the context of object manipulation. In this paper, we argue that affordance should be task-/instruction-dependent, which is overlook…