activity
20242026
collaborators
Showing 2025Show all

31 papers · 1 filter

cs.RO2025

Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives

Shuanghao Bai, Wenxuan Song, Jiayi Chen +14

Recent advances in vision, language, and multimodal learning have significantly accelerated progress in robotic foundation models, with robotic manipulation remaining one of the mo…

cs.CV2025

Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration

Yuhang Han, Xuyang Liu, Zihan Zhang +6

The quadratic complexity of Multimodal Large Language Models (MLLMs) with respect to context length poses significant computational and memory challenges, hindering their real-worl…

cs.CV2025

Rethinking Target Label Conditioning in Adversarial Attacks: A 2D Tensor-Guided Generative Approach

Hangyu Liu, Bo Peng, Pengxiang Ding +1

Compared to single-target adversarial attacks, multi-target attacks have garnered significant attention due to their ability to generate adversarial images for multiple target clas…

cs.CV2025

SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning

Yang Liu, Ming Ma, Xiaomin Yu +5

Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for inte…

cs.LG2025

Iterative Refinement of Flow Policies in Probability Space for Online Reinforcement Learning

Mingyang Sun, Pengxiang Ding, Weinan Zhang +1

While behavior cloning with flow/diffusion policies excels at learning complex skills from demonstrations, it remains vulnerable to distributional shift, and standard RL methods st…

cs.RO2025

Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model

Fuhao Li, Wenxuan Song, Han Zhao +5

Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instructions and execute precise actions. However, most VLAs are buil…