activity
20142022
most citedYOLO9000: Better, Faster, Stronger

438 citations · 612 across the 8 of their papers we have counts for

collaborators

9 papers

cs.CV202227 cited

Patching open-vocabulary models by interpolating weights

Gabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre +5

Open-vocabulary models like CLIP achieve high accuracy across many image classification tasks. However, there are still settings where their zero-shot performance is far from optim…

cs.CV2022

Break and Make: Interactive Structural Understanding Using LEGO Bricks

Aaron Walsman, Muru Zhang, Klemen Kotar +3

Visual understanding of geometric structures with complex spatial relationships is a fundamental component of human intelligence. As children, we learn how to reason about structur…

cs.CV2022

MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound

Rowan Zellers, Jiasen Lu, Ximing Lu +7

As humans, we navigate a multimodal world, building a holistic understanding from all our senses. We introduce MERLOT Reserve, a model that represents videos jointly over time -- t…

cs.CV20212 cited

Forward Compatible Training for Large-Scale Embedding Retrieval Systems

Vivek Ramanujan, Pavan Kumar Anasosalu Vasu, Ali Farhadi +2

In visual retrieval systems, updating the embedding model requires recomputing features for every piece of data. This expensive process is referred to as backfilling. Recently, the…

cs.CV2016438 cited

YOLO9000: Better, Faster, Stronger

Joseph Redmon, Ali Farhadi

We introduce YOLO9000, a state-of-the-art, real-time object detection system that can detect over 9000 object categories. First we propose various improvements to the YOLO detectio…

cs.CV20165 cited

Commonly Uncommon: Semantic Sparsity in Situation Recognition

Mark Yatskar, Vicente Ordonez, Luke Zettlemoyer +1

Semantic sparsity is a common challenge in structured visual classification problems; when the output space is complex, the vast majority of the possible predictions are rarely, if…