4 papers
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
Jinghui Lu, Jiayi Guan, Zhijian Huang +47
Chain-of-Thought (CoT) reasoning has become a powerful driver of trajectory prediction in VLA-based autonomous driving, yet its autoregressive nature imposes a latency cost that is…
VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation
Chaofan Zhang, Peng Hao, Xiaoge Cao +3
While vision-language models have advanced significantly, their application in language-conditioned robotic manipulation is still underexplored, especially for contact-rich tasks t…
BCTR: Bidirectional Conditioning Transformer for Scene Graph Generation
Peng Hao, Weilong Wang, Xiaobing Wang +5
Scene Graph Generation (SGG) remains a challenging task due to its compositional property. Previous approaches improve prediction efficiency through end-to-end learning. However, t…
TLA: Tactile-Language-Action Model for Contact-Rich Manipulation
Peng Hao, Chaofan Zhang, Dingzhe Li +4
Significant progress has been made in vision-language models. However, language-conditioned robotic manipulation for contact-rich tasks remains underexplored, particularly in terms…