149 citations · 302 across the 8 of their papers we have counts for
5 papers · 1 filter
Early Attentive Sparsification Accelerates Neural Speech Transcription
Zifei Xu, Sayeh Sharify, Hesham Mostafa +3
Transformer-based neural speech processing has achieved state-of-the-art performance. Since speech audio signals are known to be highly compressible, here we seek to accelerate neu…
On Local Aggregation in Heterophilic Graphs
Hesham Mostafa, Marcel Nassar, Somdeb Majumdar
Many recent works have studied the performance of Graph Neural Networks (GNNs) in the context of graph homophily - a label-dependent measure of connectivity. Traditional GNNs gener…
Robust Federated Learning Through Representation Matching and Adaptive Hyper-parameters
Hesham Mostafa
Federated learning is a distributed, privacy-aware learning scenario which trains a single model on data belonging to several clients. Each client trains a local model on its data…
Single-bit-per-weight deep convolutional neural networks without batch-normalization layers for embedded systems
Mark D. McDonnell, Hesham Mostafa, Runchun Wang +1
Batch-normalization (BN) layers are thought to be an integrally important layer type in today's state-of-the-art deep convolutional neural networks for computer vision tasks such a…
Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization
Hesham Mostafa, Xin Wang
Modern deep neural networks are typically highly overparameterized. Pruning techniques are able to remove a significant fraction of network parameters with little loss in accuracy.…