16 citations · 16 across the 2 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023
Video OWL-ViT: Temporally-consistent open-world localization in video
Georg Heigold, Matthias Minderer, Alexey Gritsenko +5
We present an architecture and a training recipe that adapts pre-trained open-world image models to localization in videos. Understanding the open visual world (without being const…
cs.CV2023★ 16 cited
Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution
Mostafa Dehghani, Basil Mustafa, Josip Djolonga +12
The ubiquitous and demonstrably suboptimal choice of resizing images to a fixed resolution before processing them with computer vision models has not yet been successfully challeng…