60 citations · 102 across the 12 of their papers we have counts for
16 papers
Zero-shot Image Captioning by Anchor-augmented Vision-Language Space Alignment
Junyang Wang, Yi Zhang, Ming Yan +2
CLIP (Contrastive Language-Image Pre-Training) has shown remarkable zero-shot transfer capabilities in cross-modal correlation tasks such as visual classification and image retriev…
Generating Persuasive Responses to Customer Reviews with Multi-Source Prior Knowledge in E-commerce
Bo Chen, Jiayi Liu, Mieradilijiang Maimaiti +2
Customer reviews usually contain much information about one's online shopping experience. While positive reviews are beneficial to the stores, negative ones will largely influence…
MGIMN: Multi-Grained Interactive Matching Network for Few-shot Text Classification
Jianhai Zhang, Mieradilijiang Maimaiti, Xing Gao +2
Text classification struggles to generalize to unseen classes with very few labeled text instances per class. In such a few-shot learning (FSL) setting, metric-based meta-learning…
Auto-MLM: Improved Contrastive Learning for Self-supervised Multi-lingual Knowledge Retrieval
Wenshen Xu, Mieradilijiang Maimaiti, Yuanhang Zheng +2
Contrastive learning (CL) has become a ubiquitous approach for several natural language processing (NLP) downstream tasks, especially for question answering (QA). However, the majo…
Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual Grounding
Jiabo Ye, Junfeng Tian, Ming Yan +5
Visual grounding focuses on establishing fine-grained alignment between vision and natural language, which has essential applications in multimodal reasoning systems. Existing meth…
K-AID: Enhancing Pre-trained Language Models with Domain Knowledge for Question Answering
Fu Sun, Feng-Lin Li, Ruize Wang +3
Knowledge enhanced pre-trained language models (K-PLMs) are shown to be effective for many public tasks in the literature but few of them have been successfully applied in practice…