activity
20172025
most citedSurrogate Gradient Learning in Spiking Neural Networks

149 citations · 302 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2025

Early Attentive Sparsification Accelerates Neural Speech Transcription

Zifei Xu, Sayeh Sharify, Hesham Mostafa +3

Transformer-based neural speech processing has achieved state-of-the-art performance. Since speech audio signals are known to be highly compressible, here we seek to accelerate neu…

cs.LG20211 cited

On Local Aggregation in Heterophilic Graphs

Hesham Mostafa, Marcel Nassar, Somdeb Majumdar

Many recent works have studied the performance of Graph Neural Networks (GNNs) in the context of graph homophily - a label-dependent measure of connectivity. Traditional GNNs gener…

cs.LG201921 cited

Robust Federated Learning Through Representation Matching and Adaptive Hyper-parameters

Hesham Mostafa

Federated learning is a distributed, privacy-aware learning scenario which trains a single model on data belonging to several clients. Each client trains a local model on its data…

cs.LG2019

Single-bit-per-weight deep convolutional neural networks without batch-normalization layers for embedded systems

Mark D. McDonnell, Hesham Mostafa, Runchun Wang +1

Batch-normalization (BN) layers are thought to be an integrally important layer type in today's state-of-the-art deep convolutional neural networks for computer vision tasks such a…

cs.LG2019121 cited

Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization

Hesham Mostafa, Xin Wang

Modern deep neural networks are typically highly overparameterized. Pruning techniques are able to remove a significant fraction of network parameters with little loss in accuracy.…