collaborators

6 papers

cs.CV2025

Reliable Representation Learning for Incomplete Multi-View Missing Multi-Label Classification

Chengliang Liu, Jie Wen, Yong Xu +3

As a cross-topic of multi-view learning and multi-label classification, multi-view multi-label classification has gradually gained traction in recent years. The application of mult…

cs.CV2025

Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation

Han Liu, Yinwei Wei, Fan Liu +3

Multimodal information (e.g., visual, acoustic, and textual) has been widely used to enhance representation learning for micro-video recommendation. For integrating multimodal info…

cs.CV2024

DeepFake-Adapter: Dual-Level Adapter for DeepFake Detection

Rui Shao, Tianxing Wu, Liqiang Nie +1

Existing deepfake detection methods fail to generalize well to unseen or degraded samples, which can be attributed to the over-fitting of low-level forgery patterns. Here we argue…

cs.MM2024

Self-Training Boosted Multi-Factor Matching Network for Composed Image Retrieval

Haokun Wen, Xuemeng Song, Jianhua Yin +3

The composed image retrieval (CIR) task aims to retrieve the desired target image for a given multimodal query, i.e., a reference image with its corresponding modification text. Th…

cs.CV2024

Revisiting Context Aggregation for Image Matting

Qinglin Liu, Xiaoqian Lv, Quanling Meng +5

Traditional studies emphasize the significance of context information in improving matting performance. Consequently, deep learning-based matting methods delve into designing pooli…

cs.CL2024

Multimodal Dialog Systems with Dual Knowledge-enhanced Generative Pretrained Language Model

Xiaolin Chen, Xuemeng Song, Liqiang Jing +3

Text response generation for multimodal task-oriented dialog systems, which aims to generate the proper text response given the multimodal context, is an essential yet challenging…