45 citations · 55 across the 9 of their papers we have counts for
11 papers · 1 filter
ResidualViT for Efficient Temporally Dense Video Encoding
Mattia Soldan, Fabian Caba Heilbron, Bernard Ghanem +2
Several video understanding tasks, such as natural language temporal video grounding, temporal activity localization, and audio description generation, require "temporally dense" r…
OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection
Shuming Liu, Chen Zhao, Fatimah Zohra +10
Temporal action detection (TAD) is a fundamental video understanding task that aims to identify human actions and localize their temporal boundaries in videos. Although this field…
Compressed-Language Models for Understanding Compressed File Formats: a JPEG Exploration
Juan C. Pérez, Alejandro Pardo, Mattia Soldan +3
This study investigates whether Compressed-Language Models (CLMs), i.e. language models operating on raw byte streams from Compressed File Formats~(CFFs), can understand files comp…
Towards Automated Movie Trailer Generation
Dawit Mureja Argaw, Mattia Soldan, Alejandro Pardo +4
Movie trailers are an essential tool for promoting films and attracting audiences. However, the process of creating trailers can be time-consuming and expensive. To streamline this…
Boundary-Denoising for Video Activity Localization
Mengmeng Xu, Mattia Soldan, Jialin Gao +3
Video activity localization aims at understanding the semantic content in long untrimmed videos and retrieving actions of interest. The retrieved action with its start and end loca…
Localizing Moments in Long Video Via Multimodal Guidance
Wayner Barrios, Mattia Soldan, Alberto Mario Ceballos-Arroyo +2
The recent introduction of the large-scale, long-form MAD and Ego4D datasets has enabled researchers to investigate the performance of current state-of-the-art methods for video gr…