25 citations · 30 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 25 cited
Multimodal Foundation Models: From Specialists to General-Purpose Assistants
Chunyuan Li, Zhe Gan, Zhengyuan Yang +4
This paper presents a comprehensive survey of the taxonomy and evolution of multimodal foundation models that demonstrate vision and vision-language capabilities, focusing on the t…
cs.CV2021★ 5 cited
MLP Architectures for Vision-and-Language Modeling: An Empirical Study
Yixin Nie, Linjie Li, Zhe Gan +6
We initiate the first empirical study on the use of MLP architectures for vision-and-language (VL) fusion. Through extensive experiments on 5 VL tasks and 5 robust VQA benchmarks,…