1 paper
Chao Xue, Chaofan Zhang, Wenxuan Ma +3
World action models jointly predict future visual observations and actions, whereas existing tactile-aware variants typically represent future touch as an image or latent stream wi…