11 citations · 14 across the 4 of their papers we have counts for
4 papers
ESceme: Vision-and-Language Navigation with Episodic Scene Memory
Qi Zheng, Daqing Liu, Chaoyue Wang +3
Vision-and-language navigation (VLN) simulates a visual agent that follows natural-language navigation instructions in real-world scenes. Existing approaches have made enormous pro…
Cross-Modal Contrastive Learning for Robust Reasoning in VQA
Qi Zheng, Chaoyue Wang, Daqing Liu +2
Multi-modal reasoning in visual question answering (VQA) has witnessed rapid progress recently. However, most reasoning models heavily rely on shortcuts learned from training data,…
Bypass Network for Semantics Driven Image Paragraph Captioning
Qi Zheng, Chaoyue Wang, Dadong Wang
Image paragraph captioning aims to describe a given image with a sequence of coherent sentences. Most existing methods model the coherence through the topic transition that dynamic…
Visual Superordinate Abstraction for Robust Concept Learning
Qi Zheng, Chaoyue Wang, Dadong Wang +1
Concept learning constructs visual representations that are connected to linguistic semantics, which is fundamental to vision-language tasks. Although promising progress has been m…