4 papers
Thicker and Quicker: A Jumbo Token for Fast Plain Vision Transformers
Anthony Fuller, Yousef Yassin, Daniel G. Kyrollos +2
ViTs are general and accurate, and address many tasks, but ViTs are slow, and are not always practical when efficiency is key. Existing methods for faster ViTs design hybrid non-Vi…
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
Anthony Fuller, Yousef Yassin, Junfeng Wen +4
Vision transformers are ever larger, more accurate, and more expensive to compute. The expense is even more extreme at high resolution as the number of tokens grows quadratically w…
Corporate Needs You to Find the Difference: Revisiting Submodular and Supermodular Ratio Optimization Problems
Elfarouk Harb, Yousef Yassin, Chandra Chekuri
We study the problem of minimizing or maximizing the average value of a submodular or supermodular set function over non-empty subsets $ S \s…
LookHere: Vision Transformers with Directed Attention Generalize and Extrapolate
Anthony Fuller, Daniel G. Kyrollos, Yousef Yassin +1
High-resolution images offer more information about scenes that can improve model accuracy. However, the dominant model architecture in computer vision, the vision transformer (ViT…