activity
20182022
most citedLess is More: ClipBERT for Video-and-Language Learning via Sparse Sampling

49 citations · 141 across the 7 of their papers we have counts for

collaborators

12 papers

cs.CV20221 cited

Perceiver-VL: Efficient Vision-and-Language Modeling with Iterative Latent Attention

Zineng Tang, Jaemin Cho, Jie Lei +1

We present Perceiver-VL, a vision-and-language framework that efficiently handles high-dimensional multimodal inputs such as long videos and text. Powered by the iterative latent c…

cs.CV20226 cited

LoopITR: Combining Dual and Cross Encoder Architectures for Image-Text Retrieval

Jie Lei, Xinlei Chen, Ning Zhang +4

Dual encoders and cross encoders have been widely used for image-text retrieval. Between the two, the dual encoder encodes the image and text independently followed by a dot produc…

cs.CV202133 cited

VIMPAC: Video Pre-Training via Masked Token Prediction and Contrastive Learning

Hao Tan, Jie Lei, Thomas Wolf +1

Video understanding relies on perceiving the global content and modeling its internal connections (e.g., causality, movement, and spatio-temporal correspondence). To learn these in…

cs.CV202138 cited

VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation

Linjie Li, Jie Lei, Zhe Gan +12

Most existing video-and-language (VidL) research focuses on a single dataset, or multiple datasets of a single task. In reality, a truly useful VidL system is expected to be easily…

cs.CV20212 cited

Adversarial VQA: A New Benchmark for Evaluating the Robustness of VQA Models

Linjie Li, Jie Lei, Zhe Gan +1

Benefiting from large-scale pre-training, we have witnessed significant performance boost on the popular Visual Question Answering (VQA) task. Despite rapid progress, it remains un…

cs.CV202149 cited

Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling

Jie Lei, Linjie Li, Luowei Zhou +4

The canonical approach to video-and-language learning (e.g., video question answering) dictates a neural model to learn from offline-extracted dense video features from vision mode…