1 paper
Swarnim Jain, Shangzhe Wu
Video models are increasingly used to predict what happens next in a scene, yet the metrics commonly used to compare their outputs say little about whether the predicted objects mo…