6 papers
Hi5: Synthetic Data for Inclusive, Robust, Hand Pose Estimation
Masum Hasan, Cengiz Ozel, Nina Long +5
Hand pose estimation plays a vital role in capturing subtle nonverbal cues essential for understanding human affect. However, collecting diverse, expressive real-world data remains…
WikiVideo: Article Generation from Multiple Videos
Alexander Martin, Reno Kriz, William Gantt Walden +5
We introduce the task of grounded article generation with the goal of creating a Wikipedia-style article from multiple diverse videos about real-world events -- from natural disast…
MMMORRF: Multimodal Multilingual Modularized Reciprocal Rank Fusion
Saron Samuel, Dan DeGenaro, Jimena Guallar-Blasco +13
Videos inherently contain multiple modalities, including visual events, text overlays, sounds, and speech, all of which are important for retrieval. However, state-of-the-art multi…
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
Arun Reddy, Alexander Martin, Eugene Yang +7
In this work, we tackle the problem of text-to-video retrieval (T2VR). Inspired by the success of late interaction techniques in text-document, text-image, and text-video retrieval…
MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval
Reno Kriz, Kate Sanders, David Etter +10
Efficiently retrieving and synthesizing information from large-scale multimodal collections has become a critical challenge. However, existing video retrieval datasets suffer from…
Cross-Document Event-Keyed Summarization
William Walden, Pavlo Kuchmiichuk, Alexander Martin +5
Event-keyed summarization (EKS) requires summarizing a specific event described in a document given the document text and an event representation extracted from it. In this work, w…