Learning a Text-Video Embedding from Incomplete and Heterogeneous Data
arXiv:1804.02516
Abstract
Joint understanding of video and language is an active research area with many applications. Prior work in this domain typically relies on learning text-video embeddings. One difficulty with this approach, however, is the lack of large-scale annotated video-caption datasets for training. To address this issue, we aim at learning text-video embeddings from heterogeneous data sources. To this end, we propose a Mixture-of-Embedding-Experts (MEE) model with ability to handle missing input modalities during training. As a result, our framework can learn improved text-video embeddings simultaneously from image and video datasets. We also show the generalization of MEE to other input modalities such as face descriptors. We evaluate our method on the task of video retrieval and report results for the MPII Movie Description and MSR-VTT datasets. The proposed MEE model demonstrates significant improvements and outperforms previously reported methods on both text-to-video and video-to-text retrieval tasks. Code is available at: https://github.com/antoine77340/Mixture-of-Embedding-Experts
The paper had a major update in January 2020 after a bug we found in the codebase that affected many results
References in corpus (5)
Cited by in corpus (44)
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
- Use What You Have: Video Retrieval Using Representations From Collaborative Experts
- Learning Video Representations using Contrastive Bidirectional Transformer
- Dual Encoding for Video Retrieval by Text
- Self-Supervised MultiModal Versatile Networks
- ActionCLIP: A New Paradigm for Video Action Recognition
- Self-Supervised Learning for Videos: A Survey
- MDMMT: Multidomain Multimodal Transformer for Video Retrieval
- Thinking Fast and Slow: Efficient Text-to-Visual Retrieval with Transformers
- CLIP2Video: Mastering Video-Text Retrieval via Image CLIP
- A Comprehensive Study of Deep Video Action Recognition
- Reading-strategy Inspired Visual Representation Learning for Text-to-Video Retrieval
- Audio Retrieval with Natural Language Queries: A Benchmark Study
- Focal Visual-Text Attention for Visual Question Answering
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning
- Neural ranking models for document retrieval
- Learning Spatiotemporal Features via Video and Text Pair Discrimination
- TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval
- Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers
- T2VLAD: Global-Local Sequence Alignment for Text-Video Retrieval
- Projection-Free Optimization on Uniformly Convex Sets
- The End-of-End-to-End: A Video Understanding Pentathlon Challenge (2020)
- HANet: Hierarchical Alignment Networks for Video-Text Retrieval
- Video-Text Pre-training with Learned Regions
- Semantic Role Aware Correlation Transformer for Text to Video Retrieval
- Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions
- Affine Invariant Analysis of Frank-Wolfe on Strongly Convex Sets
- An Empirical Study of Excitation and Aggregation Design Adaptions in CLIP4Clip for Video-Text Retrieval
- TEAM-Net: Multi-modal Learning for Video Action Recognition with Partial Decoding
- Rethinking movie genre classification with fine-grained semantic clustering
- Connecting Language and Vision for Natural Language-Based Vehicle Retrieval
- Video Moment Retrieval with Text Query Considering Many-to-Many Correspondence Using Potentially Relevant Pair
- Tree-Augmented Cross-Modal Encoding for Complex-Query Video Retrieval
- Rudder: A Cross Lingual Video and Text Retrieval Dataset
- Semantic Search of Memes on Twitter
- ActBERT: Learning Global-Local Video-Text Representations
- Interactive Video Retrieval with Dialog
- MTVR: Multilingual Moment Retrieval in Videos
- SwAMP: Swapped Assignment of Multi-Modal Pairs for Cross-Modal Retrieval
- Video-aided Unsupervised Grammar Induction
- E-Sports Talent Scouting Based on Multimodal Twitch Stream Data
- Retrieving and Highlighting Action with Spatiotemporal Reference
- Multi-modal Transformer for Video Retrieval
- CrossCLR: Cross-modal Contrastive Learning For Multi-modal Video Representations