22 citations · 74 across the 8 of their papers we have counts for
10 papers
Text and Patterns: For Effective Chain of Thought, It Takes Two to Tango
Aman Madaan, Amir Yazdanbakhsh
The past decade has witnessed dramatic gains in natural language processing and an unprecedented scaling of large language models. These developments have been accelerated by the a…
GRANITE: A Graph Neural Network Model for Basic Block Throughput Estimation
Ondrej Sykora, Phitchaya Mangpo Phothilimthana, Charith Mendis +1
Analytical hardware performance models yield swift estimation of desired hardware performance metrics. However, developing these analytical models for modern processors with sophis…
Training Recipe for N:M Structured Sparsity with Decaying Pruning Mask
Sheng-Chun Kao, Amir Yazdanbakhsh, Suvinay Subramanian +3
Sparsity has become one of the promising methods to compress and accelerate Deep Neural Networks (DNNs). Among different categories of sparsity, structured sparsity has gained more…
Accelerating Attention through Gradient-Based Learned Runtime Pruning
Zheng Li, Soroush Ghodrati, Amir Yazdanbakhsh +2
Self-attention is a key enabler of state-of-art accuracy for various transformer-based Natural Language Processing models. This attention mechanism calculates a correlation score f…
Rethinking Co-design of Neural Architectures and Hardware Accelerators
Yanqi Zhou, Xuanyi Dong, Berkin Akin +7
Neural architectures and hardware accelerators have been two driving forces for the progress in deep learning. Previous works typically attempt to optimize hardware given a fixed m…
Apollo: Transferable Architecture Exploration
Amir Yazdanbakhsh, Christof Angermueller, Berkin Akin +7
The looming end of Moore's Law and ascending use of deep learning drives the design of custom accelerators that are optimized for specific neural architectures. Architecture explor…