6 papers
VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer
Songqiao Hu, Zeyi Liu, Shuang Liu +5
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in generalizing across diverse robotic manipulation tasks. However, deploying these models in unstruct…
Multimodal Benchmark for Safety Assessment in Industrial Inspection Scenarios
Zeyi Liu, Shuang Liu, Jihai Min +7
With the rapid development of industrial intelligence and unmanned inspection, reliable perception and safety assessment for AI systems in complex and dynamic industrial sites has…
Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models
Yifan Ye, Jiaqi Ma, Jun Cen +1
Vision-Language-Action (VLA) models pretrained on large-scale multimodal datasets have emerged as powerful foundations for robotic perception and control. However, their massive sc…
HiMaCon: Discovering Hierarchical Manipulation Concepts from Unlabeled Multi-Modal Data
Ruizhe Liu, Pei Zhou, Qian Luo +4
Effective generalization in robotic manipulation requires representations that capture invariant patterns of interaction across environments and tasks. We present a self-supervised…
Self-evolved Imitation Learning in Simulated World
Yifan Ye, Jun Cen, Jing Chen +1
Imitation learning has been a trend recently, yet training a generalist agent across multiple tasks still requires large-scale expert demonstrations, which are costly and labor-int…
Continual Learning for Segment Anything Model Adaptation
Jinglong Yang, Yichen Wu, Jun Cen +3
Although the current different types of SAM adaptation methods have achieved promising performance for various downstream tasks, such as prompt-based ones and adapter-based ones, m…