11 papers
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
Yulu Pan, Han Yi, Seongsu Ha +4
True video intelligence demands more than recognizing what is visible: it requires reasoning about why events unfold, predicting what would change under different conditions, and d…
How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction
Sejoon Jun, Hai Nguyen-Truong, Luigi Seminara +1
Predicting how a person's first-person view will evolve (what action will follow, what plan completes a task, whether an in-progress shot will score) is fundamentally under-specifi…
RECIPE: Procedural Planning via Grounding in Instructional Video
Luigi Seminara, Antonino Furnari, Lorenzo Torresani
Visual planning asks a model to generate the remaining steps of a procedure in natural language given a partial video context and a goal. Progress on this task is bottlenecked by a…
ViewBridge: Curriculum Knowledge Distillation for Activity View-Invariance Under Extreme Viewpoint Changes
Arjun Somayazulu, Efi Mavroudi, Changan Chen +2
Traditional methods for view-invariant learning rely on controlled multi-view training data with minimal scene clutter. However, they struggle with in-the-wild videos that exhibit…
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
Shraman Pramanick, Effrosyni Mavroudi, Yale Song +3
We introduce ED-VTG, a method for fine-grained video temporal grounding utilizing multi-modal large language models. Our approach harnesses the capabilities of multimodal LLMs to j…
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Jang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi +26
Vision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The researc…