25 papers
CAC-VLA: Context-Gated Action Conditioning for Vision-Language-Action Models
Yifu Xiong, Wenhao Yu, Jiaxuan Lin +5
Vision-Language-Action (VLA) models have become a promising paradigm for generalist robot manipulation, where visual-language representations are used to condition continuous actio…
DynaHMRC: Decentralized Heterogeneous Multi-Robot Collaboration for Dynamic Tasks with Large Language Models
Wenhao Yu, Yu'ang Xie, Yifan Duan +5
Large language models (LLMs) provide robots with richer task understanding and adaptability, making them promising for coordinating heterogeneous multi-robot systems in long-horizo…
GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation
Yuan Zhang, Shiqi Zhang, Yedong Shen +11
Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot emb…
Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics
Open-H-Embodiment Consortium, :, Nigel Nelson +213
Autonomous medical robots hold promise to improve patient outcomes, reduce provider workload, democratize access to care, and enable superhuman precision. However, autonomous medic…
CORP: A Multi-Modal Dataset for Campus-Oriented Roadside Perception Tasks
Beibei Wang, Zijian Yu, Lu Zhang +8
Numerous roadside perception datasets have been introduced to propel advancements in autonomous driving and intelligent transportation systems research and development. However, it…
Needle in a Haystack: Tracking UAVs from Massive Noise in Real-World 5G-A Base Station Data
Chengzhen Meng, Chenming He, Yidong Jiang +5
The potential usage of UAVs in daily life has made monitoring them essential. However, existing systems for monitoring UAVs typically rely on cameras, LiDARs, or radars, whose limi…