From the 2 of 12 linked papers with an AI index.
12 papers
A Glimpse into Long-term Physical Coexistence with Intelligent Robots
Weiqi Jin, Peijun Tang, Kuncheng Luo +5
The paper presents PHILIA, a modular multi‑robot system that separates high‑level reasoning from low‑level robot execution via a robot‑gateway interface, enabling long‑term, person…
Towards Predictive, Aligned, and Scalable Robot Learning
Peijun Tang, Shangjin Xie, Baifu Huang +6
The paper introduces Lumo-2, a latent world-action model that reasons about future physical dynamics in a shared latent space to generate robot actions, using a multi‑stage alignme…
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
Chubin Zhang, Jianan Wang, Zifeng Gao +5
Generalist Vision-Language-Action models remain constrained by the scarcity of robotic data relative to the abundance of human video demonstrations. Existing Latent Action Models a…
SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
Peiwen Sun, Shiqiang Lang, Dongming Wu +8
With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still struggle with diverse applications su…
MetaphorVU: Towards Metaphorical Video Understanding
Zhuoqun Li, Boxi Cao, Guiping Jiang +13
Metaphorical videos are prevalent across various real-world scenarios to convey complex ideas, and understanding them typically requires high-order cognitive capabilities. The lack…
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data
Yiyang Fu, Chubin Zhang, Shukai Gong +7
It is infeasible to encompass all possible disturbances within the training dataset. This raises a critical question regarding the robustness of Vision-Language-Action (VLA) models…