2 papers
cs.CV2022
UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual Dialog
Cheng Chen, Yudong Zhu, Zhenshan Tan +4
Visual Dialog aims to answer multi-round, interactive questions based on the dialog history and image content. Existing methods either consider answer ranking and generating indivi…
cs.CV2020
Learning Dual Semantic Relations with Graph Attention for Image-Text Matching
Keyu Wen, Xiaodong Gu, Qingrong Cheng
Image-Text Matching is one major task in cross-modal information processing. The main challenge is to learn the unified visual and textual representations. Previous methods that pe…