activity
20182022
most citedVALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation

38 citations · 48 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV20223 cited

CPL: Counterfactual Prompt Learning for Vision and Language Models

Xuehai He, Diji Yang, Weixi Feng +7

Prompt tuning is a new few-shot transfer learning technique that only tunes the learnable prompt for pre-trained vision and language models such as CLIP. However, existing prompt t…

cs.CV2022

Anticipating the Unseen Discrepancy for Vision and Language Navigation

Yujie Lu, Huiliang Zhang, Ping Nie +4

Vision-Language Navigation requires the agent to follow natural language instructions to reach a specific target. The large discrepancy between seen and unseen environments makes i…

cs.CV20225 cited

Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning

Juncheng Li, Junlin Xie, Long Qian +6

Temporal grounding in videos aims to localize one target video segment that semantically corresponds to a given query sentence. Thanks to the semantic diversity of natural language…

cs.CV2021

Are Gender-Neutral Queries Really Gender-Neutral? Mitigating Gender Bias in Image Search

Jialu Wang, Yang Liu, Xin Eric Wang

Internet search affects people's cognition of the world, so mitigating biases in search results and learning fair models is imperative for social good. We study a unique gender bia…

cs.CV202138 cited

VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation

Linjie Li, Jie Lei, Zhe Gan +12

Most existing video-and-language (VidL) research focuses on a single dataset, or multiple datasets of a single task. In reality, a truly useful VidL system is expected to be easily…

cs.CV20212 cited

L2C: Describing Visual Differences Needs Semantic Understanding of Individuals

An Yan, Xin Eric Wang, Tsu-Jui Fu +1

Recent advances in language and vision push forward the research of captioning a single image to describing visual differences between image pairs. Suppose there are two images, I_…