most citedOn Latency Predictors for Neural Architecture Search

1 citations · 2 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CL2024

Radial Networks: Dynamic Layer Routing for High-Performance Large Language Models

Jordan Dotzel, Yash Akhauri, Ahmed S. AbouElhamayed +3

Large language models (LLMs) often struggle with strict memory, latency, and power demands. To meet these demands, various forms of dynamic sparsity have been proposed that reduce…

cs.LG20241 cited

Encodings for Prediction-based Neural Architecture Search

Yash Akhauri, Mohamed S. Abdelfattah

Predictor-based methods have substantially enhanced Neural Architecture Search (NAS) optimization. The efficacy of these predictors is largely influenced by the method of encoding…

cs.LG20241 cited

On Latency Predictors for Neural Architecture Search

Yash Akhauri, Mohamed S. Abdelfattah

Efficient deployment of neural networks (NN) requires the co-optimization of accuracy and latency. For example, hardware-aware neural architecture search has been used to automatic…

cs.AR2023

M4BRAM: Mixed-Precision Matrix-Matrix Multiplication in FPGA Block RAMs

Yuzong Chen, Jordan Dotzel, Mohamed S. Abdelfattah

Mixed-precision quantization is a popular approach for compressing deep neural networks (DNNs). However, it is challenging to scale the performance efficiently with mixed-precision…

cs.LG2023

DiviML: A Module-based Heuristic for Mapping Neural Networks onto Heterogeneous Platforms

Yassine Ghannane, Mohamed S. Abdelfattah

Datacenters are increasingly becoming heterogeneous, and are starting to include specialized hardware for networking, video processing, and especially deep learning. To leverage th…

cs.LG2023

Multi-Predict: Few Shot Predictors For Efficient Neural Architecture Search

Yash Akhauri, Mohamed S. Abdelfattah

Many hardware-aware neural architecture search (NAS) methods have been developed to optimize the topology of neural networks (NN) with the joint objectives of higher accuracy and l…