5 citations · 5 across the 6 of their papers we have counts for
4 papers · 1 filter
DataComp-VLM: Improved Open Datasets for Vision-Language Models
Matteo Farina, Vishaal Udandarao, Thao Nguyen +34
Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curat…
DiffCLIP: Differential Attention Meets CLIP
Hasan Abed Al Kader Hammoud, Bernard Ghanem
We propose DiffCLIP, a novel vision-language model that extends the differential attention mechanism to CLIP architectures. Differential attention was originally developed for larg…
On Pretraining Data Diversity for Self-Supervised Learning
Hasan Abed Al Kader Hammoud, Tuhin Das, Fabio Pizzati +3
We explore the impact of training with more diverse datasets, characterized by the number of unique samples, on the performance of self-supervised learning (SSL) under a fixed comp…
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
Hasan Abed Al Kader Hammoud, Hani Itani, Fabio Pizzati +3
We present SynthCLIP, a CLIP model trained on entirely synthetic text-image pairs. Leveraging recent text-to-image (TTI) networks and large language models (LLM), we generate synth…