collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2025

Test-Time Adaptation for Combating Missing Modalities in Egocentric Videos

Merey Ramazanova, Alejandro Pardo, Bernard Ghanem +1

Understanding videos that contain multiple modalities is crucial, especially in egocentric videos, where combining various sensory inputs significantly improves tasks like action r…

cs.CV2025

OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection

Shuming Liu, Chen Zhao, Fatimah Zohra +10

Temporal action detection (TAD) is a fundamental video understanding task that aims to identify human actions and localize their temporal boundaries in videos. Although this field…

cs.CV2024

MatchDiffusion: Training-free Generation of Match-cuts

Alejandro Pardo, Fabio Pizzati, Tong Zhang +4

Match-cuts are powerful cinematic tools that create seamless transitions between scenes, delivering strong visual and metaphorical connections. However, crafting match-cuts is a ch…

cs.CV2024

Generative Timelines for Instructed Visual Assembly

Alejandro Pardo, Jui-Hsien Wang, Bernard Ghanem +3

The objective of this work is to manipulate visual timelines (e.g. a video) through natural language instructions, making complex timeline editing tasks accessible to non-expert or…

cs.CV2024

Compressed-Language Models for Understanding Compressed File Formats: a JPEG Exploration

Juan C. Pérez, Alejandro Pardo, Mattia Soldan +3

This study investigates whether Compressed-Language Models (CLMs), i.e. language models operating on raw byte streams from Compressed File Formats~(CFFs), can understand files comp…

cs.CV2024

Exploring Missing Modality in Multimodal Egocentric Datasets

Merey Ramazanova, Alejandro Pardo, Humam Alwassel +1

Multimodal video understanding is crucial for analyzing egocentric videos, where integrating multiple sensory signals significantly enhances action recognition and moment localizat…