14 papers
Implicit Virtual Leader: Decentralized Vision-Only Relative Pose Estimation for Multi-Robot Formations
Shiyuan Yang, Zelin Wang, Zhijia Tao +8
Classical leader-follower formation control suffers from single points of failure and error propagation, and relies on absolute localization sensors that are ill-suited for GPS-den…
FabriVLA: A Lightweight Vision-Language-Action Model with Conformal Action Chunk Uncertainty
Shiyuan Yang, Borong Zhang, Jizheng Zhang +5
Vision-Language-Action (VLA) models have become a leading paradigm for general purpose robotic manipulation, but their computational cost and limited uncertainty awareness hinder p…
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
Borong Zhang, Jiahao Li, Jiachen Shen +8
While Vision-Language-Action models (VLAs) are rapidly advancing toward generalist robot policies, quantitatively characterizing their capability boundaries and failure modes remai…
Investigating Memory in Model-Free RL with POPGym Arcade
Zekang Wang, Zhe He, Borong Zhang +2
How should we analyze memory in deep RL? We introduce tools for analyzing policies under partial observability and revealing how agents use memory to make decisions. To utilize the…
EDT: Efficient and Effective Decision Transformer with Experience-Aware Sampling for Robotic Manipulation
Kaiyan Zhao, Borong Zhang, Yiming Wang +4
In reinforcement learning (RL) for robotic manipulation, the Decision Transformer (DT) has emerged as an effective framework for addressing long-horizon tasks. However, DT's perfor…
RedVLA: Physical Red Teaming for Vision-Language-Action Models
Yuhao Zhang, Borong Zhang, Jiaming Fan +4
The real-world deployment of Vision-Language-Action (VLA) models remains limited by the risk of unpredictable and irreversible physical harm. However, we currently lack effective m…