1 citations · 1 across the 1 of their papers we have counts for
3 papers
PaliGemma: A versatile 3B VLM for transfer
Lucas Beyer, Andreas Steiner, André Susano Pinto +32
PaliGemma is an open Vision-Language Model (VLM) that is based on the SigLIP-So400m vision encoder and the Gemma-2B language model. It is trained to be a versatile and broadly know…
Evaluating Numerical Reasoning in Text-to-Image Models
Ivana Kajić, Olivia Wiles, Isabela Albuquerque +4
Text-to-image generative models are capable of producing high-quality images that often faithfully depict concepts described using natural language. In this work, we comprehensivel…
Improving fine-grained understanding in image-text pre-training
Ioana Bica, Anastasija Ilić, Matthias Bauer +8
We introduce SPARse Fine-grained Contrastive Alignment (SPARC), a simple method for pretraining more fine-grained multimodal representations from image-text pairs. Given that multi…