activity
20182026
most citedHOP: History-and-Order Aware Pre-training for Vision-and-Language Navigation

8 citations · 10 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV20232 cited

Align before Search: Aligning Ads Image to Text for Accurate Cross-Modal Sponsored Search

Yuanmin Tang, Jing Yu, Keke Gai +4

Cross-Modal sponsored search displays multi-modal advertisements (ads) when consumers look for desired products by natural language queries in search engines. Since multi-modal ads…

cs.CV2023

Context-I2W: Mapping Images to Context-dependent Words for Accurate Zero-Shot Composed Image Retrieval

Yuanmin Tang, Jing Yu, Keke Gai +4

Different from Composed Image Retrieval task that requires expensive labels for training task-specific models, Zero-Shot Composed Image Retrieval (ZS-CIR) involves diverse tasks wi…

cs.CV20233 cited

March in Chat: Interactive Prompting for Remote Embodied Referring Expression

Yanyuan Qiao, Yuankai Qi, Zheng Yu +2

Many Vision-and-Language Navigation (VLN) tasks have been proposed in recent years, from room-based to object-based and indoor to outdoor. The REVERIE (Remote Embodied Referring Ex…

cs.CV20228 cited

HOP: History-and-Order Aware Pre-training for Vision-and-Language Navigation

Yanyuan Qiao, Yuankai Qi, Yicong Hong +3

Pre-training has been adopted in a few of recent works for Vision-and-Language Navigation (VLN). However, previous pre-training methods for VLN either lack the ability to predict f…

cs.CV20222 cited

MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question Answering

Yang Ding, Jing Yu, Bang Liu +3

Knowledge-based visual question answering requires the ability of associating external knowledge for open-ended cross-modal scene understanding. One limitation of existing solution…

cs.CV2021

Proposal-free One-stage Referring Expression via Grid-Word Cross-Attention

Wei Suo, Mengyang Sun, Peng Wang +1

Referring Expression Comprehension (REC) has become one of the most important tasks in visual reasoning, since it is an essential step for many vision-and-language tasks such as vi…