1 paper
Ayush K. Rai, Kyle Min, Tarun Krishna +3
Masked video modeling~(MVM) has emerged as a highly effective pre-training strategy for visual foundation models, whereby the model reconstructs masked spatiotemporal tokens using…