works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.RO2026

Native Video-Action Pretraining for Generalizable Robot Control

Qihang Zhang, Lin Li, Luyao Zhang +26

The paper introduces LingBot-VA 2.0, a video-action foundation model designed specifically for robot control, featuring a semantic visual-action tokenizer, causal pretraining, a sp…

cs.CV2026

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Shuailei Ma, Jiaqi Liao, Xinyang Wang +24

Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inheren…

cs.RO2026

From Foundation to Application: Improving VLA Models in Practice

Wei Wu, Fangjing Wang, Fan Lu +21

Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bri…

cs.CV2026

Vision Pretraining for Dense Spatial Perception

Zelin Fu, Bin Tan, Changjiang Sun +6

Dense spatial perception is essential for physical intelligence, where visual systems are expected to recover structured, metric, and actionable representations from pixel observat…

cs.RO2026

AetheRock: An Arm-Worn Robot Teaching System for Force-Guided Vision-Tactile Learning

Hong Li, Yue Xu, Yihan Tang +8

Force and tactile sensing are indispensable in contact-rich manipulation. However, force-aware robot learning faces critical challenges due to the incompatible assembly of tactile…

cs.CV2026

Tac-DINO: Learning Vision-Tactile Features with Patch Alignment

Hong Li, Yankang Dong, Yue Xu +8

Touch is the primary medium through which humans interact with the environment. Currently, tactile learning mainly focuses on image-level pretraining or alignment. However, tactile…