1 paper · 1 filter
Sneh Pillai
Training vision-language models for image-text alignment typically requires large datasets to achieve robust performance. In low-data scenarios, standard contrastive learning can s…