56 citations · 106 across the 23 of their papers we have counts for
4 papers · 1 filter
Learned Queries for Efficient Local Attention
Moab Arar, Ariel Shamir, Amit H. Bermano
Vision Transformers (ViT) serve as powerful vision models. Unlike convolutional neural networks, which dominated vision research in previous years, vision transformers enjoy the ab…
Rhythm is a Dancer: Music-Driven Motion Synthesis with Global Structure
Andreas Aristidou, Anastasios Yiannakidis, Kfir Aberman +3
Synthesizing human motion with a global structure, such as a choreography, is a challenging task. Existing methods tend to concentrate on local smooth pose transitions and neglect…
Ordered Attention for Coherent Visual Storytelling
Tom Braude, Idan Schwartz, Alexander Schwing +1
We address the problem of visual storytelling, i.e., generating a story for a given sequence of images. While each sentence of the story should describe a corresponding image, a co…
InAugment: Improving Classifiers via Internal Augmentation
Moab Arar, Ariel Shamir, Amit Bermano
Image augmentation techniques apply transformation functions such as rotation, shearing, or color distortion on an input image. These augmentations were proven useful in improving…