429 citations · 639 across the 30 of their papers we have counts for
16 papers · 1 filter
Multi-Dimensional Hyena for Spatial Inductive Bias
Itamar Zimerman, Lior Wolf
In recent years, Vision Transformers have attracted increasing interest from computer vision researchers. However, the advantage of these transformers over CNNs is only fully manif…
Box-based Refinement for Weakly Supervised and Unsupervised Localization Tasks
Eyal Gomel, Tal Shaharabany, Lior Wolf
It has been established that training a box-based detector network can enhance the localization performance of weakly supervised and unsupervised methods. Moreover, we extend this…
2-D SSM: A General Spatial Layer for Visual Transformers
Ethan Baron, Itamar Zimerman, Lior Wolf
A central objective in computer vision is to design models with appropriate 2-D inductive bias. Desiderata for 2D inductive bias include two-dimensional position awareness, dynamic…
AutoSAM: Adapting SAM to Medical Images by Overloading the Prompt Encoder
Tal Shaharabany, Aviad Dahan, Raja Giryes +1
The recently introduced Segment Anything Model (SAM) combines a clever architecture and large quantities of training data to obtain remarkable image segmentation capabilities. Howe…
Gradient Adjusting Networks for Domain Inversion
Erez Sheffi, Michael Rotman, Lior Wolf
StyleGAN2 was demonstrated to be a powerful image generation engine that supports semantic editing. However, in order to manipulate a real-world image, one first needs to be able t…
Zero-Shot Video Captioning with Evolving Pseudo-Tokens
Yoad Tewel, Yoav Shalev, Roy Nadler +2
We introduce a zero-shot video captioning method that employs two frozen networks: the GPT-2 language model and the CLIP image-text matching model. The matching score is used to st…