Video Highlight Prediction Using Audience Chat Reactions
arXiv:1707.08559
Abstract
Sports channel video portals offer an exciting domain for research on multimodal, multilingual analysis. We present methods addressing the problem of automatic video highlight prediction based on joint visual features and textual analysis of the real-world audience discourse with complex slang, in both English and traditional Chinese. We present a novel dataset based on League of Legends championships recorded from North American and Taiwanese Twitch.tv channels (will be released for further research), and demonstrate strong results on these using multimodal, character-level CNN-RNN model architectures.
EMNLP 2017
References in corpus (10)
- Hierarchical Question-Image Co-Attention for Visual Question Answering
- VQA: Visual Question Answering
- Dynamic Memory Networks for Visual and Textual Question Answering
- Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
- Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond
- Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
- Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research
- Visual Storytelling
- A Dataset for Movie Description
- Real-Time Video Highlights for Yahoo Esports