Sentiment-oriented Transformer-based Variational Autoencoder Network for Live Video Commenting
arXiv:2404.12782 · doi:10.1145/3633334
Abstract
Automatic live video commenting is with increasing attention due to its significance in narration generation, topic explanation, etc. However, the diverse sentiment consideration of the generated comments is missing from the current methods. Sentimental factors are critical in interactive commenting, and lack of research so far. Thus, in this paper, we propose a Sentiment-oriented Transformer-based Variational Autoencoder (So-TVAE) network which consists of a sentiment-oriented diversity encoder module and a batch attention module, to achieve diverse video commenting with multiple sentiments and multiple semantics. Specifically, our sentiment-oriented diversity encoder elegantly combines VAE and random mask mechanism to achieve semantic diversity under sentiment guidance, which is then fused with cross-modal features to generate live video comments. Furthermore, a batch attention module is also proposed in this paper to alleviate the problem of missing sentimental samples, caused by the data imbalance, which is common in live videos as the popularity of videos varies. Extensive experiments on Livebot and VideoIC datasets demonstrate that the proposed So-TVAE outperforms the state-of-the-art methods in terms of the quality and diversity of generated comments. Related code is available at https://github.com/fufy1024/So-TVAE.
27 pages, 10 figures, ACM Transactions on Multimedia Computing, Communications and Applications, 2024
References in corpus (10)
- Adam: A Method for Stochastic Optimization
- Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
- Semi-Supervised Learning with Deep Generative Models
- Improved Image Captioning via Policy Gradient optimization of SPIDEr
- Video Description: A Survey of Methods, Datasets and Evaluation Metrics
- Diverse and Accurate Image Description Using a Variational Auto-Encoder with an Additive Gaussian Encoding Space
- Topic-Guided Variational Autoencoders for Text Generation
- Herding Effect based Attention for Personalized Time-Sync Video Recommendation
- Time-sync Video Tag Extraction Using Semantic Association Graph
- BA^2M: A Batch Aware Attention Module for Image Classification