5 papers
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
Shichao Fan, Kun Wu, Zhengping Che +12
Recent progress in large-scale robotic datasets and vision-language models (VLMs) has advanced research on vision-language-action (VLA) models. However, existing VLA models still f…
Reasoning-Driven Anomaly Detection and Localization with Image-Level Supervision
Yizhou Jin, Yuezhu Feng, Jinjin Zhang +3
Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning and perceptual abilities for anomaly detection. However, most approaches remain confined to…
LIBERO-X: Robustness Litmus for Vision-Language-Action Models
Guodong Wang, Chenkai Zhang, Qingjie Liu +4
Reliable benchmarking is critical for advancing Vision-Language-Action (VLA) models, as it reveals their generalization, robustness, and alignment of perception with language-drive…
Diffusion Trajectory-guided Policy for Long-horizon Robot Manipulation
Shichao Fan, Quantao Yang, Yajie Liu +4
Recently, Vision-Language-Action models (VLA) have advanced robot imitation learning, but high data collection costs and limited demonstrations hinder generalization and current im…
Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation
Yongchao Feng, Yajie Liu, Shuai Yang +13
Vision-Language Model (VLM) have gained widespread adoption in Open-Vocabulary (OV) object detection and segmentation tasks. Despite they have shown promise on OV-related tasks, th…