5 papers
ALPHA-PIM: Analysis of Linear Algebraic Processing for High-Performance Graph Applications on a Real Processing-In-Memory System
Marzieh Barkhordar, Alireza Tabatabaeian, Mohammad Sadrosadati +5
Processing large-scale graph datasets is computationally intensive and time-consuming. Processor-centric CPU and GPU architectures, commonly used for graph applications, often face…
Faster Vertex Cover Algorithms on GPUs with Component-Aware Parallel Branching
Hussein Amro, Basel Fakhri, Amer E. Mouawad +1
Algorithms for finding minimum or bounded vertex covers in graphs use a branch-and-reduce strategy, which involves exploring a highly imbalanced search tree. Prior GPU solutions as…
PruneX: A Hierarchical Communication-Efficient System for Distributed CNN Training with Structured Pruning
Alireza Olama, Andreas Lundell, Izzat El Hajj +2
Inter-node communication bandwidth increasingly constrains distributed training at scale on multi-node GPU clusters. While compact models are the ultimate deployment target, conven…
Low-latency control system for feedback experiments with optical tweezer arrays
Amir H. Dadpour, Timur Khayrullin, Fouad Afiouni +4
We present and characterize a modular, open-source system to perform feedback control experiments on configurations of atoms and molecules in arrays of optical tweezers. The system…
Efficient algorithms to solve atom reconfiguration problems. III. The bird and batching algorithms and other parallel implementations on GPUs
Fouad Afiouni, Remy El Sabeh, Naomi Nishimura +3
We present efficient implementations of atom reconfiguration algorithms for both CPUs and GPUs, along with a batching routine to merge displacement operations for parallel executio…