2 papers
cs.CV2026
RigidBench: Evaluating Rigid-Body Physics in Video Generation Models
Swarnim Jain, Shangzhe Wu
Video models are increasingly used to predict what happens next in a scene, yet the metrics commonly used to compare their outputs say little about whether the predicted objects mo…
cs.CV2026
HiResNets: Native Full-HD Video Recognition with Foveal Residual Streams
Shivani Mall, Swarnim Jain, Joao F. Henriques
Much of the recent progress in image and video recognition has come at the cost of memory: larger models, increased resolution, and longer temporal contexts. An inevitable componen…