activity
20202024
most citedAre we done with ImageNet?

74 citations · 89 across the 5 of their papers we have counts for

collaborators

10 papers

cs.CV20241 cited

Context-Aware Multimodal Pretraining

Karsten Roth, Zeynep Akata, Dima Damen +2

Large-scale multimodal representation learning successfully optimizes for zero-shot transfer at test time. Yet the standard pretraining paradigm (contrastive learning on large amou…

cs.CV2024

Active Data Curation Effectively Distills Large-Scale Multimodal Models

Vishaal Udandarao, Nikhil Parthasarathy, Muhammad Ferjad Naeem +6

Knowledge distillation (KD) is the de facto standard for compressing large-scale models into smaller ones. Prior works have explored ever more complex KD strategies involving diffe…

cs.CV2024

A Practitioner's Guide to Continual Multimodal Pretraining

Karsten Roth, Vishaal Udandarao, Sebastian Dziadzio +7

Multimodal foundation models serve numerous applications at the intersection of vision and language. Still, despite being pretrained on extensive data, they become outdated over ti…

cs.LG2024

Reflecting on the State of Rehearsal-free Continual Learning with Pretrained Models

Lukas Thede, Karsten Roth, Olivier J. Hénaff +2

With the advent and recent ubiquity of foundation models, continual learning (CL) has recently shifted from continual training from scratch to the continual adaptation of pretraine…

cs.CV2024

Memory Consolidation Enables Long-Context Video Understanding

Ivana Balažević, Yuge Shi, Pinelopi Papalampidi +3

Most transformer-based video encoders are limited to short temporal contexts due to their quadratic complexity. While various attempts have been made to extend this context, this h…

cs.LG2023

Fantastic Gains and Where to Find Them: On the Existence and Prospect of General Knowledge Transfer between Any Pretrained Model

Karsten Roth, Lukas Thede, Almut Sophia Koepke +3

Training deep networks requires various design decisions regarding for instance their architecture, data augmentation, or optimization. In this work, we find these training variati…