11 papers
Uncovering Linguistic Fragility in Vision-Language-Action Models via Diversity-Aware Red Teaming
Baoshun Tong, Haoran He, Ling Pan +2
Vision-Language-Action (VLA) models have achieved remarkable success in robotic manipulation. However, their robustness to linguistic nuances remains a critical, under-explored saf…
Referring-Aware Visuomotor Policy Learning for Closed-Loop Manipulation
Jiahua Ma, Yiran Qin, Xin Wen +5
This paper addresses a fundamental problem of visuomotor policy learning for robotic manipulation: how to enhance robustness in out-of-distribution execution errors or dynamically…
DDP-WM: Disentangled Dynamics Prediction for Efficient World Models
Shicheng Yin, Kaixuan Yin, Weixing Chen +3
World models are essential for autonomous robotic planning. However, the substantial computational overhead of existing dense Transformerbased models significantly hinders real-tim…
AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning
Weixing Chen, Dafeng Chi, Yang Liu +7
The automated generation of layouts is vital for embodied intelligence and autonomous systems, supporting applications from virtual environment construction to home robot deploymen…
DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Foundation Models
Shicheng Yin, Kaixuan Yin, Yang Liu +2
The content-agnostic, fixed-grid tokenizers used by standard large-scale vision models like Vision Transformer (ViT) and Vision Mamba (Vim) represent a fundamental performance bott…
3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussians
Zeming Wei, Junyi Lin, Yang Liu +4
3D affordance reasoning is essential in associating human instructions with the functional regions of 3D objects, facilitating precise, task-oriented manipulations in embodied AI.…