1 citations · 1 across the 2 of their papers we have counts for
3 papers
Vision-Language Model Dialog Games for Self-Improvement
Ksenia Konyushkova, Christos Kaplanis, Serkan Cabi +1
The increasing demand for high-quality, diverse training data poses a significant bottleneck in advancing vision-language models (VLMs). This paper presents VLM Dialog Games, a nov…
Synth: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings
Sahand Sharifzadeh, Christos Kaplanis, Shreya Pathak +5
The creation of high-quality human-labeled image-caption datasets presents a significant bottleneck in the development of Visual-Language Models (VLMs). In this work, we investigat…
Improving fine-grained understanding in image-text pre-training
Ioana Bica, Anastasija Ilić, Matthias Bauer +8
We introduce SPARse Fine-grained Contrastive Alignment (SPARC), a simple method for pretraining more fine-grained multimodal representations from image-text pairs. Given that multi…