Unsupervised Semantic Parsing of Video Collections
arXiv:1506.08438
Abstract
Human communication typically has an underlying structure. This is reflected in the fact that in many user generated videos, a starting point, ending, and certain objective steps between these two can be identified. In this paper, we propose a method for parsing a video into such semantic steps in an unsupervised way. The proposed method is capable of providing a semantic "storyline" of the video composed of its objective steps. We accomplish this using both visual and language cues in a joint generative model. The proposed method can also provide a textual description for each of the identified semantic steps and video segments. We evaluate this method on a large number of complex YouTube videos and show results of unprecedented quality for this intricate and impactful problem.
References in corpus (3)
Cited by in corpus (7)
- Tripping through time: Efficient Localization of Activities in Videos
- Connectionist Temporal Modeling for Weakly Supervised Action Labeling
- VirtualHome: Simulating Household Activities via Programs
- RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes
- Unsupervised Visual-Linguistic Reference Resolution in Instructional Videos
- Joint Discovery of Object States and Manipulation Actions
- Discriminatively Learned Hierarchical Rank Pooling Networks