1 paper
Xizhou Bu, Jiexi Lyu, Fulei Sun +3
Learning latent actions from large-scale videos is crucial for the pre-training of scalable embodied foundation models, yet existing methods often struggle with action-irrelevant d…