10 papers
RoboVista: Evaluating Vision Language Models for Diverse Robot Applications
Shuangyu Xie, Kaiyuan Chen, Ziyang Chen +8
Diverse applications for robotics, such as industry and agriculture, require robots to operate across various embodiments, changing visual conditions, and complex planning. Vision-…
KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning
Yixuan Huang, Bowen Li, Vaibhav Saxena +9
Robotic systems that interact with the physical world must reason about kinematic and dynamic constraints imposed by their own embodiment, their environment, and the task at hand.…
Learning Adaptive Reasoning Paths for Efficient Visual Reasoning
Yixu Huang, Tinghui Zhu, Muhao Chen
Visual reasoning models (VRMs) have recently shown strong cross-modal reasoning capabilities by integrating visual perception with language reasoning. However, they often suffer fr…
Hierarchical DLO Routing with Reinforcement Learning and In-Context Vision-language Models
Mingen Li, Houjian Yu, Yixuan Huang +3
Long-horizon routing tasks of deformable linear objects (DLOs), such as cables and ropes, are common in industrial assembly lines and everyday life. These tasks are particularly ch…
LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer
Lihan Zha, Asher J. Hancock, Mingtong Zhang +5
A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodim…
Composable Visual Tokenizers with Generator-Free Diagnostics of Learnability
Bingchen Zhao, Qiushan Guo, Ye Wang +3
We introduce CompTok, a training framework for learning visual tokenizers whose tokens are enhanced for compositionality. CompTok uses a token-conditioned diffusion decoder. By emp…