1 paper
Utkarsh A. Mishra, Yongxin Chen, Danfei Xu +3
Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these priors after finetuning on ro…