31 papers · 1 filter
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
Shuanghao Bai, Wenxuan Song, Jiayi Chen +14
Recent advances in vision, language, and multimodal learning have significantly accelerated progress in robotic foundation models, with robotic manipulation remaining one of the mo…
Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration
Yuhang Han, Xuyang Liu, Zihan Zhang +6
The quadratic complexity of Multimodal Large Language Models (MLLMs) with respect to context length poses significant computational and memory challenges, hindering their real-worl…
Rethinking Target Label Conditioning in Adversarial Attacks: A 2D Tensor-Guided Generative Approach
Hangyu Liu, Bo Peng, Pengxiang Ding +1
Compared to single-target adversarial attacks, multi-target attacks have garnered significant attention due to their ability to generate adversarial images for multiple target clas…
SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
Yang Liu, Ming Ma, Xiaomin Yu +5
Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for inte…
Iterative Refinement of Flow Policies in Probability Space for Online Reinforcement Learning
Mingyang Sun, Pengxiang Ding, Weinan Zhang +1
While behavior cloning with flow/diffusion policies excels at learning complex skills from demonstrations, it remains vulnerable to distributional shift, and standard RL methods st…
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
Fuhao Li, Wenxuan Song, Han Zhao +5
Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instructions and execute precise actions. However, most VLAs are buil…