Video2GIF: Automatic Generation of Animated GIFs from Video
arXiv:1605.04850
Abstract
We introduce the novel problem of automatically generating animated GIFs from video. GIFs are short looping video with no sound, and a perfect combination between image and video that really capture our attention. GIFs tell a story, express emotion, turn events into humorous moments, and are the new wave of photojournalism. We pose the question: Can we automate the entirely manual and elaborate process of GIF creation by leveraging the plethora of user generated GIF content? We propose a Robust Deep RankNet that, given a video, generates a ranked list of its segments according to their suitability as GIF. We train our model to learn what visual content is often selected for GIFs by using over 100K user generated GIFs and their corresponding video sources. We effectively deal with the noisy web data by proposing a novel adaptive Huber loss in the ranking formulation. We show that our approach is robust to outliers and picks up several patterns that are frequently present in popular animated GIFs. On our new large-scale benchmark dataset, we show the advantage of our approach over several state-of-the-art methods.
Accepted to CVPR 2016
References in corpus (2)
Cited by in corpus (16)
- TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering
- Polysemous Visual-Semantic Embedding for Cross-Modal Retrieval
- Real-Time Video Highlights for Yahoo Esports
- To Click or Not To Click: Automatic Selection of Beautiful Thumbnails from Videos
- Cross-Modal Retrieval with Implicit Concept Association
- A Deep Ranking Model for Spatio-Temporal Highlight Detection from a 360 Video
- ElasticPlay: Interactive Video Summarization with Dynamic Time Budgets
- Less is More: Learning Highlight Detection from Video Duration
- Sentence Guided Temporal Modulation for Dynamic Video Thumbnail Generation
- Vis-DSS: An Open-Source toolkit for Visual Data Selection and Summarization
- A Memory Network Approach for Story-based Temporal Summarization of 360° Videos
- Contextually Customized Video Summaries via Natural Language
- Detecting Cognitive Appraisals from Facial Expressions for Interest Recognition
- Summarizing First-Person Videos from Third Persons' Points of Views
- A Mobile Robot Generating Video Summaries of Seniors' Indoor Activities
- Demystifying Multi-Faceted Video Summarization: Tradeoff Between Diversity,Representation, Coverage and Importance