papers

Publications (28)

cs.CL2024

Towards Foundation Models for Knowledge Graph Reasoning

Mikhail Galkin, Xinyu Yuan, Hesham Mostafa +2

Foundation models in language and vision have the ability to run inference on any textual and visual inputs thanks to the transferable representations such as a vocabulary of token…

cs.NE2020

Synaptic Plasticity Dynamics for Deep Continuous Local Learning (DECOLLE)

Jacques Kaiser, Hesham Mostafa, Emre Neftci

A growing body of work underlines striking similarities between biological neural networks and recurrent, binary neural networks. A relatively smaller body of work, however, discus…

cs.CV2018

NullHop: A Flexible Convolutional Neural Network Accelerator Based on Sparse Representations of Feature Maps

Alessandro Aimar, Hesham Mostafa, Enrico Calabrese +8

Convolutional neural networks (CNNs) have become the dominant neural network architecture for solving many state-of-the-art (SOA) visual processing tasks. Even though Graphical Pro…

cs.OH2024

PARSAC: Fast, Human-quality Floorplanning for Modern SoCs with Complex Design Constraints

Hesham Mostafa, Uday Mallappa, Mikhail Galkin +2

The floorplanning of Systems-on-a-Chip (SoCs) and of chip sub-systems is a crucial step in the physical design flow as it determines the optimal shapes and locations of the blocks…

cs.LG2024

Distributed Training of Large Graph Neural Networks with Variable Communication Rates

Juan Cervino, Md Asadullah Turja, Hesham Mostafa +2

Training Graph Neural Networks (GNNs) on large graphs presents unique challenges due to the large memory and computing requirements. Distributed GNN training, where the graph is pa…

cs.LG2022

Sequential Aggregation and Rematerialization: Distributed Full-batch Training of Graph Neural Networks on Large Graphs

Hesham Mostafa

We present the Sequential Aggregation and Rematerialization (SAR) scheme for distributed full-batch training of Graph Neural Networks (GNNs) on large graphs. Large-scale training o…

cs.AI2020

Permutohedral-GCN: Graph Convolutional Networks with Global Attention

Hesham Mostafa, Marcel Nassar

Graph convolutional networks (GCNs) update a node's feature vector by aggregating features from its neighbors in the graph. This ignores potentially useful contributions from dista…

cs.CV2026

DiRotQ: Rotation-Aware Quantization for 4-bit Diffusion Transformers

Sayeh Sharify, Mahsa Salmani, Hesham Mostafa

Diffusion Transformers (DiTs) achieve state-of-the-art image generation quality but incur substantial memory and computational costs at inference. While aggressive Post-Training Qu…

cs.NE2018

A learning framework for winner-take-all networks with stochastic synapses

Hesham Mostafa, Gert Cauwenberghs

Many recent generative models make use of neural networks to transform the probability distribution of a simple low-dimensional noise process into the complex distribution of the d…

cs.LG2021

On Local Aggregation in Heterophilic Graphs

Hesham Mostafa, Marcel Nassar, Somdeb Majumdar

Many recent works have studied the performance of Graph Neural Networks (GNNs) in the context of graph homophily - a label-dependent measure of connectivity. Traditional GNNs gener…

cs.LG2025

Early Attentive Sparsification Accelerates Neural Speech Transcription

Zifei Xu, Sayeh Sharify, Hesham Mostafa +3

Transformer-based neural speech processing has achieved state-of-the-art performance. Since speech audio signals are known to be highly compressible, here we seek to accelerate neu…

cs.LG2025

Fully-inductive Node Classification on Arbitrary Graphs

Jianan Zhao, Zhaocheng Zhu, Mikhail Galkin +3

One fundamental challenge in graph machine learning is generalizing to new graphs. Many existing methods following the inductive setup can generalize to test graphs with new struct…

cs.CV2022

Exploiting Long-Term Dependencies for Generating Dynamic Scene Graphs

Shengyu Feng, Subarna Tripathi, Hesham Mostafa +2

Dynamic scene graph generation from a video is challenging due to the temporal dynamics of the scene and the inherent temporal fluctuations of predictions. We hypothesize that capt…

cs.LG2019

Single-bit-per-weight deep convolutional neural networks without batch-normalization layers for embedded systems

Mark D. McDonnell, Hesham Mostafa, Runchun Wang +1

Batch-normalization (BN) layers are thought to be an integrally important layer type in today's state-of-the-art deep convolutional neural networks for computer vision tasks such a…

cs.NE2017

Supervised learning based on temporal coding in spiking neural networks

Hesham Mostafa

Gradient descent training techniques are remarkably successful in training analog-valued artificial neural networks (ANNs). Such training techniques, however, do not transfer easil…

cs.CV2020

Attention-based Image Upsampling

Souvik Kundu, Hesham Mostafa, Sharath Nittur Sridhar +1

Convolutional layers are an integral part of many deep neural network solutions in computer vision. Recent work shows that replacing the standard convolution operation with mechani…

cs.DC2023

FastSample: Accelerating Distributed Graph Neural Network Training for Billion-Scale Graphs

Hesham Mostafa, Adam Grabowski, Md Asadullah Turja +3

Training Graph Neural Networks(GNNs) on a large monolithic graph presents unique challenges as the graph cannot fit within a single machine and it cannot be decomposed into smaller…

cs.LG2019

Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization

Hesham Mostafa, Xin Wang

Modern deep neural networks are typically highly overparameterized. Pruning techniques are able to remove a significant fraction of network parameters with little loss in accuracy.…

q-bio.NC2015

Rhythmic inhibition allows neural networks to search for maximally consistent states

Hesham Mostafa, Lorenz K. Muller, Giacomo Indiveri

Gamma-band rhythmic inhibition is a ubiquitous phenomenon in neural circuits yet its computational role still remains elusive. We show that a model of Gamma-band rhythmic inhibitio…

cs.NE2015

Stochastic Interpretation of Quasi-periodic Event-based Systems

Hesham Mostafa, Giacomo Indiveri

Many networks used in machine learning and as models of biological neural networks make use of stochastic neurons or neuron-like units. We show that stochastic artificial neurons c…

cs.AR2024

FloorSet -- a VLSI Floorplanning Dataset with Design Constraints of Real-World SoCs

Uday Mallappa, Hesham Mostafa, Mikhail Galkin +2

Floorplanning for systems-on-a-chip (SoCs) and its sub-systems is a crucial and non-trivial step of the physical design flow. It represents a difficult combinatorial optimization p…

cs.DC2015

An event-based architecture for solving constraint satisfaction problems

Hesham Mostafa, Lorenz K. Müller, Giacomo Indiveri

Constraint satisfaction problems (CSPs) are typically solved using conventional von Neumann computing architectures. However, these architectures do not reflect the distributed nat…

cs.LG2026

MF-QAT: Multi-Format Quantization-Aware Training for Elastic Inference

Zifei Xu, Sayeh Sharify, Hesham Mostafa

Quantization-aware training (QAT) is typically performed for a single target numeric format, while practical deployments often need to choose numerical precision at inference time…

cs.LG2021

Implicit SVD for Graph Representation Learning

Sami Abu-El-Haija, Hesham Mostafa, Marcel Nassar +3

Recent improvements in the performance of state-of-the-art (SOTA) methods for Graph Representational Learning (GRL) have come at the cost of significant computational resource requ…

cs.NE2017

Hardware-efficient on-line learning through pipelined truncated-error backpropagation in binary-state networks

Hesham Mostafa, Bruno Pedroni, Sadique Sheik +1

Artificial neural networks (ANNs) trained using backpropagation are powerful learning architectures that have achieved state-of-the-art performance in various benchmarks. Significa…

cs.NE2019

Surrogate Gradient Learning in Spiking Neural Networks

Emre O. Neftci, Hesham Mostafa, Friedemann Zenke

Spiking neural networks are nature's versatile solution to fault-tolerant and energy efficient signal processing. To translate these benefits into hardware, a growing number of neu…

cs.LG2019

Robust Federated Learning Through Representation Matching and Adaptive Hyper-parameters

Hesham Mostafa

Federated learning is a distributed, privacy-aware learning scenario which trains a single model on data belonging to several clients. Each client trains a local model on its data…

cs.NE2017

Deep supervised learning using local errors

Hesham Mostafa, Vishwajith Ramesh, Gert Cauwenberghs

Error backpropagation is a highly effective mechanism for learning high-quality hierarchical features in deep networks. Updating the features or weights in one layer, however, requ…