39 citations · 72 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 26 cited
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Xi Chen, Xiao Wang, Lucas Beyer +16
This paper presents PaLI-3, a smaller, faster, and stronger vision language model (VLM) that compares favorably to similar models that are 10x larger. As part of arriving at this s…
cs.CV2023★ 39 cited
PaLI-X: On Scaling up a Multilingual Vision and Language Model
Xi Chen, Josip Djolonga, Piotr Padlewski +40
We present the training recipe and results of scaling up PaLI-X, a multilingual vision and language model, both in terms of size of the components and the breadth of its training t…
cs.CL2016★ 7 cited
Understanding Image and Text Simultaneously: a Dual Vision-Language Machine Comprehension Task
Nan Ding, Sebastian Goodman, Fei Sha +1
We introduce a new multi-modal task for computer systems, posed as a combined vision-language comprehension challenge: identifying the most suitable text describing a scene, given…