1 paper · 1 filter
Alvi Md Ishmam, Christopher Thomas
In recent years there has been enormous interest in vision-language models trained using self-supervised objectives. However, the use of large-scale datasets scraped from the web f…