5 papers
TrackMAE: Video Representation Learning via Track Mask and Predict
Renaud Vandeghen, Fida Mohammad Thoker, Marc Van Droogenbroeck +1
Masked video modeling (MVM) has emerged as a simple and scalable self-supervised pretraining paradigm, but only encodes motion information implicitly, limiting the encoding of temp…
TAPS: Task Aware Proposal Distributions for Speculative Sampling
Mohamad Zbib, Mohamad Bazzi, Ammar Mohanna +2
Speculative decoding accelerates autoregressive generation by letting a lightweight draft model propose future tokens that a larger target model then verifies in parallel. In pract…
Transformers from Compressed Representations
Juan C. Leon Alcazar, Mattia Soldan, Mohammad Saatialsoruji +4
Compressed file formats are the corner stone of efficient data storage and transmission, yet their potential for representation learning remains largely underexplored. We introduce…
ResidualViT for Efficient Temporally Dense Video Encoding
Mattia Soldan, Fabian Caba Heilbron, Bernard Ghanem +2
Several video understanding tasks, such as natural language temporal video grounding, temporal activity localization, and audio description generation, require "temporally dense" r…
Generative Timelines for Instructed Visual Assembly
Alejandro Pardo, Jui-Hsien Wang, Bernard Ghanem +3
The objective of this work is to manipulate visual timelines (e.g. a video) through natural language instructions, making complex timeline editing tasks accessible to non-expert or…