871 citations · 1.4k across the 7 of their papers we have counts for
10 papers
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Wenliang Dai, Junnan Li, Dongxu Li +6
Large-scale pre-training and instruction tuning have been successful at creating general-purpose language models with broad competence. However, building general-purpose vision-lan…
The Devil in Linear Transformer
Zhen Qin, XiaoDong Han, Weixuan Sun +4
Linear transformers aim to reduce the quadratic space-time complexity of vanilla transformers. However, they usually suffer from degraded performances on various tasks and corpus.…
LAVIS: A Library for Language-Vision Intelligence
Dongxu Li, Junnan Li, Hung Le +3
We introduce LAVIS, an open-source deep learning library for LAnguage-VISion research and applications. LAVIS aims to serve as a one-stop comprehensive library that brings recent a…
cosFormer: Rethinking Softmax in Attention
Zhen Qin, Weixuan Sun, Hui Deng +6
Transformer has shown great successes in natural language processing, computer vision, and audio processing. As one of its core components, the softmax attention helps to capture l…
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Junnan Li, Dongxu Li, Caiming Xiong +1
Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based t…
ARVo: Learning All-Range Volumetric Correspondence for Video Deblurring
Dongxu Li, Chenchen Xu, Kaihao Zhang +5
Video deblurring models exploit consecutive frames to remove blurs from camera shakes and object motions. In order to utilize neighboring sharp patches, typical methods rely mainly…