6 papers · 1 filter
Test-Time Adaptation for Combating Missing Modalities in Egocentric Videos
Merey Ramazanova, Alejandro Pardo, Bernard Ghanem +1
Understanding videos that contain multiple modalities is crucial, especially in egocentric videos, where combining various sensory inputs significantly improves tasks like action r…
OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection
Shuming Liu, Chen Zhao, Fatimah Zohra +10
Temporal action detection (TAD) is a fundamental video understanding task that aims to identify human actions and localize their temporal boundaries in videos. Although this field…
MatchDiffusion: Training-free Generation of Match-cuts
Alejandro Pardo, Fabio Pizzati, Tong Zhang +4
Match-cuts are powerful cinematic tools that create seamless transitions between scenes, delivering strong visual and metaphorical connections. However, crafting match-cuts is a ch…
Generative Timelines for Instructed Visual Assembly
Alejandro Pardo, Jui-Hsien Wang, Bernard Ghanem +3
The objective of this work is to manipulate visual timelines (e.g. a video) through natural language instructions, making complex timeline editing tasks accessible to non-expert or…
Compressed-Language Models for Understanding Compressed File Formats: a JPEG Exploration
Juan C. Pérez, Alejandro Pardo, Mattia Soldan +3
This study investigates whether Compressed-Language Models (CLMs), i.e. language models operating on raw byte streams from Compressed File Formats~(CFFs), can understand files comp…
Exploring Missing Modality in Multimodal Egocentric Datasets
Merey Ramazanova, Alejandro Pardo, Humam Alwassel +1
Multimodal video understanding is crucial for analyzing egocentric videos, where integrating multiple sensory signals significantly enhances action recognition and moment localizat…