1 citations · 2 across the 7 of their papers we have counts for
7 papers
Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models
Jordan Dotzel, Yash Akhauri, Ahmed S. AbouElhamayed +3
Large language models (LLMs) often struggle with strict memory, latency, and power demands. To meet these demands, various forms of dynamic sparsity have been proposed that reduce…
Encodings for Prediction-based Neural Architecture Search
Yash Akhauri, Mohamed S. Abdelfattah
Predictor-based methods have substantially enhanced Neural Architecture Search (NAS) optimization. The efficacy of these predictors is largely influenced by the method of encoding…
On Latency Predictors for Neural Architecture Search
Yash Akhauri, Mohamed S. Abdelfattah
Efficient deployment of neural networks (NN) requires the co-optimization of accuracy and latency. For example, hardware-aware neural architecture search has been used to automatic…
M4BRAM: Mixed-Precision Matrix-Matrix Multiplication in FPGA Block RAMs
Yuzong Chen, Jordan Dotzel, Mohamed S. Abdelfattah
Mixed-precision quantization is a popular approach for compressing deep neural networks (DNNs). However, it is challenging to scale the performance efficiently with mixed-precision…
DiviML: A Module-based Heuristic for Mapping Neural Networks onto Heterogeneous Platforms
Yassine Ghannane, Mohamed S. Abdelfattah
Datacenters are increasingly becoming heterogeneous, and are starting to include specialized hardware for networking, video processing, and especially deep learning. To leverage th…
Multi-Predict: Few Shot Predictors For Efficient Neural Architecture Search
Yash Akhauri, Mohamed S. Abdelfattah
Many hardware-aware neural architecture search (NAS) methods have been developed to optimize the topology of neural networks (NN) with the joint objectives of higher accuracy and l…