1 paper · 1 filter
Manish Kumar Govind, Dominick Reilly, Pu Wang +1
Latent action representations learned from unlabeled videos have recently emerged as a promising paradigm for pretraining vision-language-action (VLA) models without explicit robot…