4 papers
Smaller, Faster, Cheaper: Architectural Designs for Efficient Machine Learning
Steven Walton
Major advancements in the capabilities of computer vision models have been primarily fueled by rapid expansion of datasets, model parameters, and computational budgets, leading to…
Distilling Normalizing Flows
Steven Walton, Valeriy Klyukin, Maksim Artemev +3
Explicit density learners are becoming an increasingly popular technique for generative models because of their ability to better model probability distributions. They have advanta…
Efficient Image Generation with Variadic Attention Heads
Steven Walton, Ali Hassani, Xingqian Xu +2
While the integration of transformers in vision models have yielded significant improvements on vision tasks they still require significant amounts of computation for both training…
Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
Ali Hassani, Fengzhe Zhou, Aditya Kane +13
Many sparse attention mechanisms such as Neighborhood Attention have typically failed to consistently deliver speedup over the self attention baseline. This is largely due to the l…