activity
20242026
collaborators

11 papers

cs.CV2026

SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence

Yulu Pan, Han Yi, Seongsu Ha +4

True video intelligence demands more than recognizing what is visible: it requires reasoning about why events unfold, predicting what would change under different conditions, and d…

cs.CV2026

How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction

Sejoon Jun, Hai Nguyen-Truong, Luigi Seminara +1

Predicting how a person's first-person view will evolve (what action will follow, what plan completes a task, whether an in-progress shot will score) is fundamentally under-specifi…

cs.CV2026

RECIPE: Procedural Planning via Grounding in Instructional Video

Luigi Seminara, Antonino Furnari, Lorenzo Torresani

Visual planning asks a model to generate the remaining steps of a procedure in natural language given a partial video context and a goal. Progress on this task is bottlenecked by a…

cs.CV2026

ViewBridge: Curriculum Knowledge Distillation for Activity View-Invariance Under Extreme Viewpoint Changes

Arjun Somayazulu, Efi Mavroudi, Changan Chen +2

Traditional methods for view-invariant learning rely on controlled multi-view training data with minimal scene clutter. However, they struggle with in-the-wild videos that exhibit…

cs.CV2025

Enrich and Detect: Video Temporal Grounding with Multimodal LLMs

Shraman Pramanick, Effrosyni Mavroudi, Yale Song +3

We introduce ED-VTG, a method for fine-grained video temporal grounding utilizing multi-modal large language models. Our approach harnesses the capabilities of multimodal LLMs to j…

cs.CV2025

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Jang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi +26

Vision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The researc…