most citedVALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation

38 citations · 48 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV20226 cited

LoopITR: Combining Dual and Cross Encoder Architectures for Image-Text Retrieval

Jie Lei, Xinlei Chen, Ning Zhang +4

Dual encoders and cross encoders have been widely used for image-text retrieval. Between the two, the dual encoder encodes the image and text independently followed by a dot produc…

cs.CV20221 cited

Unsupervised Vision-and-Language Pre-training via Retrieval-based Multi-Granular Alignment

Mingyang Zhou, Licheng Yu, Amanpreet Singh +3

Vision-and-Language (V+L) pre-training models have achieved tremendous success in recent years on various multi-modal benchmarks. However, the majority of existing models require p…

cs.CV20223 cited

CommerceMM: Large-Scale Commerce MultiModal Representation Learning with Omni Retrieval

Licheng Yu, Jun Chen, Animesh Sinha +4

We introduce CommerceMM - a multimodal model capable of providing a diverse and granular understanding of commerce topics associated to the given piece of content (image, text, ima…

cs.CV202138 cited

VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation

Linjie Li, Jie Lei, Zhe Gan +12

Most existing video-and-language (VidL) research focuses on a single dataset, or multiple datasets of a single task. In reality, a truly useful VidL system is expected to be easily…

cs.CV2021

Connecting What to Say With Where to Look by Modeling Human Attention Traces

Zihang Meng, Licheng Yu, Ning Zhang +4

We introduce a unified framework to jointly model images, text, and human attention traces. Our work is built on top of the recent Localized Narratives annotation framework [30], w…