3 citations · 8 across the 3 of their papers we have counts for
4 papers
Efficient Zero-shot Visual Search via Target and Context-aware Transformer
Zhiwei Ding, Xuezhe Ren, Erwan David +3
Visual search is a ubiquitous challenge in natural vision, including daily tasks such as finding a friend in a crowd or searching for a car in a parking lot. Human rely heavily on…
On the Efficacy of Co-Attention Transformer Layers in Visual Question Answering
Ankur Sikarwar, Gabriel Kreiman
In recent years, multi-modal transformers have shown significant progress in Vision-Language tasks, such as Visual Question Answering (VQA), outperforming previous architectures by…
Visual Search Asymmetry: Deep Nets and Humans Share Similar Inherent Biases
Shashi Kant Gupta, Mengmi Zhang, Chia-Chien Wu +2
Visual search is a ubiquitous and often challenging daily task, exemplified by looking for the car keys at home or a friend in a crowd. An intriguing property of some classical sea…
When Pigs Fly: Contextual Reasoning in Synthetic and Natural Scenes
Philipp Bomatter, Mengmi Zhang, Dimitar Karev +3
Context is of fundamental importance to both human and machine vision; e.g., an object in the air is more likely to be an airplane than a pig. The rich notion of context incorporat…