4 papers
-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
Xiaowei Cai, Yunuo Cai, Bingao Chen +36
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-acti…
-WM: A Unified Video-Action World Model for Robotic Manipulation
Pengfei Zhou, Shengcong Chen, Di Chen +17
Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execution. We present -World…
ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
Yifan Li, Yingda Yin, Lingting Zhu +4
Reasoning-centric video object segmentation is an inherently complex task: the query often refers to dynamics, causality, and temporal interactions, rather than static appearances.…
Unified Lexical Representation for Interpretable Visual-Language Alignment
Yifan Li, Yikai Wang, Yanwei Fu +3
Visual-Language Alignment (VLA) has gained a lot of attention since CLIP's groundbreaking work. Although CLIP performs well, the typical direct latent feature alignment lacks clari…