14.1k citations · 17.1k across the 31 of their papers we have counts for
13 papers · 1 filter
A Practitioner's Guide to Continual Multimodal Pretraining
Karsten Roth, Vishaal Udandarao, Sebastian Dziadzio +7
Multimodal foundation models serve numerous applications at the intersection of vision and language. Still, despite being pretrained on extensive data, they become outdated over ti…
Non-isotropy Regularization for Proxy-based Deep Metric Learning
Karsten Roth, Oriol Vinyals, Zeynep Akata
Deep Metric Learning (DML) aims to learn representation spaces on which semantic relations can simply be expressed through predefined distance metrics. Best performing approaches c…
Integrating Language Guidance into Vision-based Deep Metric Learning
Karsten Roth, Oriol Vinyals, Zeynep Akata
Deep Metric Learning (DML) proposes to learn metric spaces which encode semantic similarities as embedding space distances. These spaces should be transferable to classes beyond th…
Multimodal Few-Shot Learning with Frozen Language Models
Maria Tsimpoukelli, Jacob Menick, Serkan Cabi +3
When trained at sufficient scale, auto-regressive language models exhibit the notable ability to learn a new language task after being prompted with just a few examples. Here, we p…
Efficient Visual Pretraining with Contrastive Detection
Olivier J. Hénaff, Skanda Koppula, Jean-Baptiste Alayrac +3
Self-supervised pretraining has been shown to yield powerful representations for transfer learning. These performance gains come at a large computational cost however, with state-o…
Perceiver: General Perception with Iterative Attention
Andrew Jaegle, Felix Gimeno, Andrew Brock +3
Biological systems perceive the world by simultaneously processing high-dimensional inputs from modalities as diverse as vision, audition, touch, proprioception, etc. The perceptio…