128 citations · 168 across the 9 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023★ 2 cited
Vision Transformers with Mixed-Resolution Tokenization
Tomer Ronen, Omer Levy, Avram Golbert
Vision Transformer models process input images by dividing them into a spatially regular grid of equal-size patches. Conversely, Transformers were originally introduced over natura…
cs.CV2023
X&Fuse: Fusing Visual Information in Text-to-Image Generation
Yuval Kirstain, Omer Levy, Adam Polyak
We introduce X&Fuse, a general approach for conditioning on visual information when generating images from text. We demonstrate the potential of X&Fuse in three different text-to-i…