17 citations · 17 across the 1 of their papers we have counts for
3 papers
cs.RO2026
In-Context World Modeling for Robotic Control
Siyin Wang, Junhao Shi, Senyu Fei +4
Modern Vision-Language-Action (VLA) models often fail to generalize to novel setups, such as altered camera viewpoints or robot morphologies, because they are typically conditioned…
cs.RO2026
Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data
Linqi Yin, Shiduo Zhang, Shenling Qiu +11
Vision-language models (VLMs) are powerful general-purpose reasoners, yet converting them into robot control policies (VLAs) is surprisingly difficult. The root cause is a two-fold…
cs.LG2021★ 17 cited
Semi-Supervised Multi-Modal Multi-Instance Multi-Label Deep Network with Optimal Transport
Yang Yang, Zhao-Yang Fu, De-Chuan Zhan +2
Complex objects are usually with multiple labels, and can be represented by multiple modal representations, e.g., the complex articles contain text and image information as well as…