5 citations · 6 across the 4 of their papers we have counts for
4 papers
A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
Benno Krojer, Mojtaba Komeili, Candace Ross +4
Existing benchmarks for assessing the spatio-temporal understanding and reasoning abilities of video language models are susceptible to score inflation due to the presence of short…
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Mido Assran, Adrien Bardes, David Fan +27
A major challenge for modern AI is to learn to understand the world and learn to act largely by observation. This paper explores a self-supervised approach that combines internet-s…
Seeing Voices: Generating A-Roll Video from Audio with Mirage
Aditi Sundararaman, Amogh Adishesha, Andrew Jaegle +10
From professional filmmaking to user-generated content, creators and consumers have long recognized that the power of video depends on the harmonious integration of what we hear (t…
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning
Han Lin, Tushar Nagarajan, Nicolas Ballas +4
Procedural video representation learning is an active research area where the objective is to learn an agent which can anticipate and forecast the future given the present video in…