activity
20242026
collaborators

10 papers

cs.RO2026

RoboVista: Evaluating Vision Language Models for Diverse Robot Applications

Shuangyu Xie, Kaiyuan Chen, Ziyang Chen +8

Diverse applications for robotics, such as industry and agriculture, require robots to operate across various embodiments, changing visual conditions, and complex planning. Vision-…

cs.RO2026

KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning

Yixuan Huang, Bowen Li, Vaibhav Saxena +9

Robotic systems that interact with the physical world must reason about kinematic and dynamic constraints imposed by their own embodiment, their environment, and the task at hand.…

cs.CV2026

Learning Adaptive Reasoning Paths for Efficient Visual Reasoning

Yixu Huang, Tinghui Zhu, Muhao Chen

Visual reasoning models (VRMs) have recently shown strong cross-modal reasoning capabilities by integrating visual perception with language reasoning. However, they often suffer fr…

cs.RO2026

Hierarchical DLO Routing with Reinforcement Learning and In-Context Vision-language Models

Mingen Li, Houjian Yu, Yixuan Huang +3

Long-horizon routing tasks of deformable linear objects (DLOs), such as cables and ropes, are common in industrial assembly lines and everyday life. These tasks are particularly ch…

cs.RO2026

LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer

Lihan Zha, Asher J. Hancock, Mingtong Zhang +5

A long-standing goal in robotics is a generalist policy that can be deployed zero-shot on new robot embodiments without per-embodiment adaptation. Despite large-scale multi-embodim…

cs.CV2026

Composable Visual Tokenizers with Generator-Free Diagnostics of Learnability

Bingchen Zhao, Qiushan Guo, Ye Wang +3

We introduce CompTok, a training framework for learning visual tokenizers whose tokens are enhanced for compositionality. CompTok uses a token-conditioned diffusion decoder. By emp…