activity
20192022
most citedOmniVL:One Foundation Model for Image-Language and Video-Language Tasks

69 citations · 178 across the 6 of their papers we have counts for

collaborators

7 papers

cs.CV202269 cited

OmniVL:One Foundation Model for Image-Language and Video-Language Tasks

Junke Wang, Dongdong Chen, Zuxuan Wu +7

This paper presents OmniVL, a new foundation model to support both image-language and video-language tasks using one universal architecture. It adopts a unified transformer-based v…

cs.CV20228 cited

Multimodal Adaptive Distillation for Leveraging Unimodal Encoders for Vision-Language Tasks

Zhecan Wang, Noel Codella, Yen-Chun Chen +8

Cross-modal encoders for vision-language (VL) tasks are often pretrained with carefully curated vision-language datasets. While these datasets reach an order of 10 million samples,…

cs.CV202138 cited

VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation

Linjie Li, Jie Lei, Zhe Gan +12

Most existing video-and-language (VidL) research focuses on a single dataset, or multiple datasets of a single task. In reality, a truly useful VidL system is expected to be easily…

cs.CV20217 cited

CUPID: Adaptive Curation of Pre-training Data for Video-and-Language Representation Learning

Luowei Zhou, Jingjing Liu, Yu Cheng +2

This work concerns video-language pre-training and representation learning. In this now ubiquitous training scheme, a model first performs pre-training on paired videos and text (e…

cs.CV20217 cited

UC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-training

Mingyang Zhou, Luowei Zhou, Shuohang Wang +4

Vision-and-language pre-training has achieved impressive success in learning multimodal representations between vision and language. To generalize this success to non-English langu…

cs.CV202149 cited

Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling

Jie Lei, Linjie Li, Luowei Zhou +4

The canonical approach to video-and-language learning (e.g., video question answering) dictates a neural model to learn from offline-extracted dense video features from vision mode…