collaborators

5 papers

cs.DC2026

ALPHA-PIM: Analysis of Linear Algebraic Processing for High-Performance Graph Applications on a Real Processing-In-Memory System

Marzieh Barkhordar, Alireza Tabatabaeian, Mohammad Sadrosadati +5

Processing large-scale graph datasets is computationally intensive and time-consuming. Processor-centric CPU and GPU architectures, commonly used for graph applications, often face…

cs.DC2025

Faster Vertex Cover Algorithms on GPUs with Component-Aware Parallel Branching

Hussein Amro, Basel Fakhri, Amer E. Mouawad +1

Algorithms for finding minimum or bounded vertex covers in graphs use a branch-and-reduce strategy, which involves exploring a highly imbalanced search tree. Prior GPU solutions as…

cs.DC2025

PruneX: A Hierarchical Communication-Efficient System for Distributed CNN Training with Structured Pruning

Alireza Olama, Andreas Lundell, Izzat El Hajj +2

Inter-node communication bandwidth increasingly constrains distributed training at scale on multi-node GPU clusters. While compact models are the ultimate deployment target, conven…

quant-ph2025

Low-latency control system for feedback experiments with optical tweezer arrays

Amir H. Dadpour, Timur Khayrullin, Fouad Afiouni +4

We present and characterize a modular, open-source system to perform feedback control experiments on configurations of atoms and molecules in arrays of optical tweezers. The system…

quant-ph2025

Efficient algorithms to solve atom reconfiguration problems. III. The bird and batching algorithms and other parallel implementations on GPUs

Fouad Afiouni, Remy El Sabeh, Naomi Nishimura +3

We present efficient implementations of atom reconfiguration algorithms for both CPUs and GPUs, along with a batching routine to merge displacement operations for parallel executio…