Video Summarization with Long Short-term Memory
arXiv:1605.08110
Abstract
We propose a novel supervised learning technique for summarizing videos by automatically selecting keyframes or key subshots. Casting the problem as a structured prediction problem on sequential data, our main idea is to use Long Short-Term Memory (LSTM), a special type of recurrent neural networks to model the variable-range dependencies entailed in the task of video summarization. Our learning models attain the state-of-the-art results on two benchmark video datasets. Detailed analysis justifies the design of the models. In particular, we show that it is crucial to take into consideration the sequential structures in videos and model them. Besides advances in modeling techniques, we introduce techniques to address the need of a large number of annotated data for training complex learning models. There, our main idea is to exploit the existence of auxiliary annotated video datasets, albeit heterogeneous in visual styles and contents. Specifically, we show domain adaptation techniques can improve summarization by reducing the discrepancies in statistical properties across those datasets.
To appear in ECCV 2016
References in corpus (4)
Cited by in corpus (27)
- Reconstructive Sequence-Graph Network for Video Summarization
- End-to-End Dense Video Captioning with Masked Transformer
- Query-Conditioned Three-Player Adversarial Network for Video Summarization
- Making a long story short: A Multi-Importance fast-forwarding egocentric videos with the emphasis on relevant objects
- Real-Time Video Highlights for Yahoo Esports
- Temporal Dynamic Graph LSTM for Action-driven Video Object Detection
- Deep 360 Pilot: Learning a Deep Agent for Piloting through 360° Sports Video
- Collaborative Summarization of Topic-Related Videos
- Learning to score the figure skating sports videos
- A Unified Framework for Generic, Query-Focused, Privacy Preserving and Update Summarization using Submodular Information Measures
- Temporal Tessellation: A Unified Approach for Video Analysis
- Estimating Blink Probability for Highlight Detection in Figure Skating Videos
- How Local is the Local Diversity? Reinforcing Sequential Determinantal Point Processes with Dynamic Ground Sets for Supervised Video Summarization
- Improving Sequential Determinantal Point Processes for Supervised Video Summarization
- Iterative Projection and Matching: Finding Structure-preserving Representatives and Its Application to Computer Vision
- A Memory Network Approach for Story-based Temporal Summarization of 360° Videos
- I Have Seen Enough: A Teacher Student Network for Video Classification Using Fewer Frames
- Automatic Comic Generation with Stylistic Multi-page Layouts and Emotion-driven Text Balloon Generation
- GPT2MVS: Generative Pre-trained Transformer-2 for Multi-modal Video Summarization
- Hierarchical Hidden Markov Jump Processes for Cancer Screening Modeling
- What I See Is What You See: Joint Attention Learning for First and Third Person Video Co-analysis
- Summarizing First-Person Videos from Third Persons' Points of Views
- From Thumbnails to Summaries - A single Deep Neural Network to Rule Them All
- Multi-View Surveillance Video Summarization via Joint Embedding and Sparse Optimization
- A Mobile Robot Generating Video Summaries of Seniors' Indoor Activities
- FFNet: Video Fast-Forwarding via Reinforcement Learning
- Customizing First Person Image Through Desired Actions