3 papers
cs.CV2021
WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
Yuqi Huo, Manli Zhang, Guangzhen Liu +32
Multi-modal pre-training models have been intensively explored to bridge vision and language in recent years. However, most of them explicitly model the cross-modal interaction bet…
cs.LG2020
AdarGCN: Adaptive Aggregation GCN for Few-Shot Learning
Jianhong Zhang, Manli Zhang, Zhiwu Lu +2
Existing few-shot learning (FSL) methods assume that there exist sufficient training samples from source classes for knowledge transfer to target classes with few training samples.…
cs.CV2018
Recursive Visual Attention in Visual Dialog
Yulei Niu, Hanwang Zhang, Manli Zhang +3
Visual dialog is a challenging vision-language task, which requires the agent to answer multi-round questions about an image. It typically needs to address two major problems: (1)…