1 paper
Junyuan Xiao, Dingkang Liang, Xin Zhou +9
Emerging multi-modal world models attempt to jointly generate videos across diverse modalities (e.g., RGB, depth, and mask), yet they fail to fully exploit the rich priors of exist…