Multi-queue Momentum Contrast for Microvideo-Product Retrieval
arXiv:2212.11471 · doi:10.1145/3539597.3570405
Abstract
The booming development and huge market of micro-videos bring new e-commerce channels for merchants. Currently, more micro-video publishers prefer to embed relevant ads into their micro-videos, which not only provides them with business income but helps the audiences to discover their interesting products. However, due to the micro-video recording by unprofessional equipment, involving various topics and including multiple modalities, it is challenging to locate the products related to micro-videos efficiently, appropriately, and accurately. We formulate the microvideo-product retrieval task, which is the first attempt to explore the retrieval between the multi-modal and multi-modal instances. A novel approach named Multi-Queue Momentum Contrast (MQMC) network is proposed for bidirectional retrieval, consisting of the uni-modal feature and multi-modal instance representation learning. Moreover, a discriminative selection strategy with a multi-queue is used to distinguish the importance of different negatives based on their categories. We collect two large-scale microvideo-product datasets (MVS and MVS-large) for evaluation and manually construct the hierarchical category ontology, which covers sundry products in daily life. Extensive experiments show that MQMC outperforms the state-of-the-art baselines. Our replication package (including code, dataset, etc.) is publicly available at https://github.com/duyali2000/MQMC.
Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining (WSDM '23), February 27-March 3, 2023, Singapore, Singapore
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Bootstrap your own latent: A new approach to self-supervised Learning
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- Contrastive Meta Learning with Behavior Multiplicity for Recommendation
- Counterfactual Reasoning for Out-of-distribution Multimodal Sentiment Analysis
- On the User Behavior Leakage from Recommender System Exposure
- Micro-video Tagging via Jointly Modeling Social Influence and Tag Relation