Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
arXiv:1506.06724
Abstract
Books are a rich source of both fine-grained information, how a character, an object or a scene looks like, as well as high-level semantics, what someone is thinking, feeling and how these states evolve through a story. This paper aims to align books to their movie releases in order to provide rich descriptive explanations for visual content that go semantically far beyond the captions available in current datasets. To align movies and books we exploit a neural sentence embedding that is trained in an unsupervised way from a large corpus of books, as well as a video-text neural embedding for computing similarities between movie clips and sentences in the book. We propose a context-aware CNN to combine information from multiple sources. We demonstrate good quantitative performance for movie/book alignment and show several qualitative examples that showcase the diversity of tasks our model can be used for.
References in corpus (7)
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Going Deeper with Convolutions
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Explain Images with Multimodal Recurrent Neural Networks
- Show and Tell: A Neural Image Caption Generator
- Efficient Structured Prediction with Latent Variables for General Graphical Models
- A Dataset for Movie Description
Cited by in corpus (81)
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation
- Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation
- BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model
- COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining
- Adversarial Feature Matching for Text Generation
- Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
- Character-level and Multi-channel Convolutional Neural Networks for Large-scale Authorship Attribution
- Language Models are Open Knowledge Graphs
- Efficient Vector Representation for Documents through Corruption
- See, Hear, and Read: Deep Aligned Representations
- Linguistic Features for Readability Assessment
- Sample Efficient Text Summarization Using a Single Pre-Trained Transformer
- HellaSwag: Can a Machine Really Finish Your Sentence?
- Efficient Adaptation of Pretrained Transformers for Abstractive Summarization
- What Would Elsa Do? Freezing Layers During Transformer Fine-Tuning
- UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning
- TwinBERT: Distilling Knowledge to Twin-Structured BERT Models for Efficient Retrieval
- Optimal Subarchitecture Extraction For BERT
- SIMILAR: Submodular Information Measures Based Active Learning In Realistic Scenarios
- What's in the Box? A Preliminary Analysis of Undesirable Content in the Common Crawl Corpus
- Evolution of transfer learning in natural language processing
- Pre-Trained Models: Past, Present and Future
- Empathetic BERT2BERT Conversational Model: Learning Arabic Language Generation with Little Data
- BERT-based Ensembles for Modeling Disclosure and Support in Conversational Social Media Text
- Accenture at CheckThat! 2020: If you say so: Post-hoc fact-checking of claims using transformer-based models
- MC-BERT: Efficient Language Pre-Training via a Meta Controller
- An End-to-End Approach to Natural Language Object Retrieval via Context-Aware Deep Reinforcement Learning
- Generating (Factual?) Narrative Summaries of RCTs: Experiments with Neural Multi-Document Summarization
- Real-Time Execution of Large-scale Language Models on Mobile
- Combining Pre-trained Word Embeddings and Linguistic Features for Sequential Metaphor Identification
- An Automated, End-to-End Framework for Modeling Attacks From Vulnerability Descriptions
- LazyFormer: Self Attention with Lazy Update
- Variance-reduced Language Pretraining via a Mask Proposal Network
- Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded Supervision
- Transformers with Competitive Ensembles of Independent Mechanisms
- WebRED: Effective Pretraining And Finetuning For Relation Extraction On The Web
- Support-Set Based Cross-Supervision for Video Grounding
- Stable, Fast and Accurate: Kernelized Attention with Relative Positional Encoding
- Unsupervised Vision-and-Language Pre-training Without Parallel Images and Captions
- Summarize, Outline, and Elaborate: Long-Text Generation via Hierarchical Supervision from Extractive Summaries
- Back-Translated Task Adaptive Pretraining: Improving Accuracy and Robustness on Text Classification
- Fine-tuning of Pre-trained Transformers for Hate, Offensive, and Profane Content Detection in English and Marathi
- Eigen Analysis of Self-Attention and its Reconstruction from Partial Computation
- Pretraining with Contrastive Sentence Objectives Improves Discourse Performance of Language Models
- Sentiment analysis in tweets: an assessment study from classical to modern text representation models
- Pay Attention when Required
- How Vulnerable Are Automatic Fake News Detection Methods to Adversarial Attacks?
- LinCE: A Centralized Benchmark for Linguistic Code-switching Evaluation
- Pun Generation with Surprise
- Explicit Pairwise Word Interaction Modeling Improves Pretrained Transformers for English Semantic Similarity Tasks
- Operationalizing a National Digital Library: The Case for a Norwegian Transformer Model
- Extract then Distill: Efficient and Effective Task-Agnostic BERT Distillation
- Undivided Attention: Are Intermediate Layers Necessary for BERT?
- Automatic learner summary assessment for reading comprehension
- Shatter: An Efficient Transformer Encoder with Single-Headed Self-Attention and Relative Sequence Partitioning
- WikiCheck: An end-to-end open source Automatic Fact-Checking API based on Wikipedia
- Robustness of on-device Models: Adversarial Attack to Deep Learning Models on Android Apps
- Bilingual Language Modeling, A transfer learning technique for Roman Urdu
- Towards Robust Pattern Recognition: A Review
- Discourse Marker Augmented Network with Reinforcement Learning for Natural Language Inference
- Plan ahead: Self-Supervised Text Planning for Paragraph Completion Task
- Semantic Text-to-Face GAN -ST^2FG
- Contrastive Learning of Visual-Semantic Embeddings
- Evaluating Document Coherence Modelling
- Semantic Frame Forecast
- LV-BERT: Exploiting Layer Variety for BERT
- Recognising Biomedical Names: Challenges and Solutions
- CALM: Continuous Adaptive Learning for Language Modeling
- Learning Joint Representations of Videos and Sentences with Web Image Search
- Evidence-Aware Inferential Text Generation with Vector Quantised Variational AutoEncoder
- Have Attention Heads in BERT Learned Constituency Grammar?
- Improving Distributed Representations of Tweets - Present and Future
- PARADE: A New Dataset for Paraphrase Identification Requiring Computer Science Domain Knowledge
- mAnI: Movie Amalgamation using Neural Imitation
- Weakly Supervised Learning of Heterogeneous Concepts in Videos
- Data-efficient Alignment of Multimodal Sequences by Aligning Gradient Updates and Internal Feature Distributions
- A Temporal Variational Model for Story Generation