295 citations · 422 across the 8 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023
Three Towers: Flexible Contrastive Learning with Pretrained Image Models
Jannik Kossen, Mark Collier, Basil Mustafa +7
We introduce Three Towers (3T), a flexible method to improve the contrastive learning of vision-language models by incorporating pretrained image classifiers. While contrastive mod…
cs.CV2021★ 29 cited
Scaling Vision with Sparse Mixture of Experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa +5
Sparsely-gated Mixture of Experts networks (MoEs) have demonstrated excellent scalability in Natural Language Processing. In Computer Vision, however, almost all performant network…