8.1k citations · 9.1k across the 8 of their papers we have counts for
4 papers · 1 filter
Video OWL-ViT: Temporally-consistent open-world localization in video
Georg Heigold, Matthias Minderer, Alexey Gritsenko +5
We present an architecture and a training recipe that adapts pre-trained open-world image models to localization in videos. Understanding the open visual world (without being const…
Sensitivity of Slot-Based Object-Centric Models to their Number of Slots
Roland S. Zimmermann, Sjoerd van Steenkiste, Mehdi S. M. Sajjadi +2
Self-supervised methods for learning object-centric representations have recently been applied successfully to various datasets. This progress is largely fueled by slot-based metho…
AudioSlots: A slot-centric generative model for audio separation
Pradyumna Reddy, Scott Wisdom, Klaus Greff +2
In a range of recent works, object-centric architectures have been shown to be suitable for unsupervised scene decomposition in the vision domain. Inspired by these methods we pres…
Scaling Vision Transformers to 22 Billion Parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa +39
The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Visio…