2 papers
cs.RO2026
VISTA: Visually Inferred Spatial ConTact Attention for Contact-Rich Manipulation
Jiayi Chen, Wenlong Dong, Yan Huang +5
Contact-rich manipulation requires precise interaction feedback. While vision-centric imitation learning is prevalent, external visual observations provide indirect and ambiguous c…
cs.RO2026
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
Haoran Yuan, Weigang Yi, Zhenyu Zhang +9
Video-Action Models (VAMs) have emerged as a promising framework for embodied intelligence, learning implicit world dynamics from raw video streams to produce temporally consistent…