6 papers
RoboVista: Evaluating Vision Language Models for Diverse Robot Applications
Shuangyu Xie, Kaiyuan Chen, Ziyang Chen +8
Diverse applications for robotics, such as industry and agriculture, require robots to operate across various embodiments, changing visual conditions, and complex planning. Vision-…
T-Rex: Tactile-Reactive Dexterous Manipulation
Dantong Niu, Zhuoyang Liu, Zekai Wang +31
The ability to react dynamically to tactile signals has long been considered crucial to agile human-level dexterity. Yet contemporary learning-based Vision-Language-Action (VLA) mo…
OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
Huang Huang, Fangchen Liu, Letian Fu +5
Vision-Language-Action (VLA) models aim to predict robotic actions based on visual observations and language instructions. Existing approaches require fine-tuning pre-trained visio…
Rethinking Patch Dependence for Masked Autoencoders
Letian Fu, Long Lian, Renhao Wang +6
In this work, we examine the impact of inter-patch dependencies in the decoder of masked autoencoders (MAE) on representation learning. We decompose the decoding mechanism for mask…
DiffusionSeeder: Seeding Motion Optimization with Diffusion for Rapid Motion Planning
Huang Huang, Balakumar Sundaralingam, Arsalan Mousavian +3
Running optimization across many parallel seeds leveraging GPU compute have relaxed the need for a good initialization, but this can fail if the problem is highly non-convex as all…
Self-Supervised Learning of Dynamic Planar Manipulation of Free-End Cables
Jonathan Wang, Huang Huang, Vincent Lim +5
Dynamic manipulation of free-end cables has applications for cable management in homes, warehouses and manufacturing plants. We present a supervised learning approach for dynamic m…