most citedUC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-training

7 citations · 20 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

GloSplat: Joint Pose-Appearance Optimization for Faster and More Accurate 3D Reconstruction

Tianyu Xiong, Rui Li, Linjie Li +1

Feature extraction, matching, structure from motion (SfM), and novel view synthesis (NVS) have traditionally been treated as separate problems with independent optimization objecti…

cs.CV2021★ 5 cited

MLP Architectures for Vision-and-Language Modeling: An Empirical Study

Yixin Nie, Linjie Li, Zhe Gan +6

We initiate the first empirical study on the use of MLP architectures for vision-and-language (VL) fusion. Through extensive experiments on 5 VL tasks and 5 robust VQA benchmarks,…

cs.CV2021

SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning

Kevin Lin, Linjie Li, Chung-Ching Lin +5

The canonical approach to video captioning dictates a caption generation model to learn from offline-extracted dense video features. These feature extractors usually operate on vid…

cs.CV2021

VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Tsu-Jui Fu, Linjie Li, Zhe Gan +4

A great challenge in video-language (VidL) modeling lies in the disconnection between fixed video representations extracted from image/video understanding models and downstream Vid…

cs.CV2021★ 2 cited

Adversarial VQA: A New Benchmark for Evaluating the Robustness of VQA Models

Linjie Li, Jie Lei, Zhe Gan +1

Benefiting from large-scale pre-training, we have witnessed significant performance boost on the popular Visual Question Answering (VQA) task. Despite rapid progress, it remains un…

cs.CV2021

Playing Lottery Tickets with Vision and Language

Zhe Gan, Yen-Chun Chen, Linjie Li +6

Large-scale pre-training has recently revolutionized vision-and-language (VL) research. Models such as LXMERT and UNITER have significantly lifted the state of the art over a wide…