3 papers
cs.RO2026
Guava: An Effective and Universal Harness for Embodied Manipulation
Haowen Liu, Xirui Li, Shaoxiong Yao +5
Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through embodied tools use offers a promising…
cs.RO2026
SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models
Haowen Liu, Shaoxiong Yao, Haonan Chen +4
Vision-Language Models (VLMs) exhibit remarkable common-sense and semantic reasoning capabilities. However, they lack a grounded understanding of physical dynamics. This limitation…
cs.RO2026
Implicit State Estimation via Video Replanning
Po-Chen Ko, Jiayuan Mao, Yu-Hsiang Fu +5
Video-based representations have gained prominence in planning and decision-making due to their ability to encode rich spatiotemporal dynamics and geometric relationships. These re…