6 papers
Reliable Representation Learning for Incomplete Multi-View Missing Multi-Label Classification
Chengliang Liu, Jie Wen, Yong Xu +3
As a cross-topic of multi-view learning and multi-label classification, multi-view multi-label classification has gradually gained traction in recent years. The application of mult…
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
Han Liu, Yinwei Wei, Fan Liu +3
Multimodal information (e.g., visual, acoustic, and textual) has been widely used to enhance representation learning for micro-video recommendation. For integrating multimodal info…
DeepFake-Adapter: Dual-Level Adapter for DeepFake Detection
Rui Shao, Tianxing Wu, Liqiang Nie +1
Existing deepfake detection methods fail to generalize well to unseen or degraded samples, which can be attributed to the over-fitting of low-level forgery patterns. Here we argue…
Self-Training Boosted Multi-Factor Matching Network for Composed Image Retrieval
Haokun Wen, Xuemeng Song, Jianhua Yin +3
The composed image retrieval (CIR) task aims to retrieve the desired target image for a given multimodal query, i.e., a reference image with its corresponding modification text. Th…
Revisiting Context Aggregation for Image Matting
Qinglin Liu, Xiaoqian Lv, Quanling Meng +5
Traditional studies emphasize the significance of context information in improving matting performance. Consequently, deep learning-based matting methods delve into designing pooli…
Multimodal Dialog Systems with Dual Knowledge-enhanced Generative Pretrained Language Model
Xiaolin Chen, Xuemeng Song, Liqiang Jing +3
Text response generation for multimodal task-oriented dialog systems, which aims to generate the proper text response given the multimodal context, is an essential yet challenging…