5 citations · 15 across the 4 of their papers we have counts for
12 papers · 1 filter
Image interpretation by iterative bottom-up top-down processing
Shimon Ullman, Liav Assif, Alona Strugatski +4
Scene understanding requires the extraction and representation of scene components together with their properties and inter-relations. We describe a model in which meaningful scene…
Detector-Free Weakly Supervised Grounding by Separation
Assaf Arbelle, Sivan Doveh, Amit Alfassy +14
Nowadays, there is an abundance of data involving images and surrounding free-form text weakly corresponding to those images. Weakly Supervised phrase-Grounding (WSG) deals with th…
What can human minimal videos tell us about dynamic recognition models?
Guy Ben-Yosef, Gabriel Kreiman, Shimon Ullman
In human vision objects and their parts can be visually recognized from purely spatial or purely temporal information but the mechanisms integrating space and time are poorly under…
Efficient Coarse-to-Fine Non-Local Module for the Detection of Small Objects
Hila Levi, Shimon Ullman
An image is not just a collection of objects, but rather a graph where each object is related to other objects through spatial and semantic relations. Using relational reasoning mo…
VQA with no questions-answers training
Ben-Zion Vatashsky, Shimon Ullman
Methods for teaching machines to answer visual questions have made significant progress in recent years, but current methods still lack important human capabilities, including inte…
Understand, Compose and Respond - Answering Visual Questions by a Composition of Abstract Procedures
Ben Zion Vatashsky, Shimon Ullman
An image related question defines a specific visual task that is required in order to produce an appropriate answer. The answer may depend on a minor detail in the image and requir…