3 papers
cs.RO2026
CoFreeVLA: Collision-Free Dual-Arm Manipulation via Vision-Language-Action Model and Risk Estimation
Xuanran Zhai, Binkai Ou, Qiaojun Yu +2
Vision Language Action (VLA) models enable instruction following manipulation, yet dualarm deployment remains unsafe due to under modeled selfcollisions between arms and grasped ob…
cs.RO2026
Spiking Neural-Invariant Kalman Fusion for Accurate Localization Using Low-Cost IMUs
Yaohua Liu, Qiao Xu, Binkai Ou
Low-cost inertial measurement units (IMUs) are widely utilized in mobile robot localization due to their affordability and ease of integration. However, their complex, nonlinear, a…
cs.MM2026
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
Qihao Zhao, Yunqi Cao, Yangyu Huang +4
Despite recent advances in multimodal large language models (MLLMs), their ability to understand and interact with music remains limited. Music understanding requires grounded reas…