7 papers
GreenGNN: Energy-Aware Windowed Communication Optimization for Distributed GNN Training
Arefin Niam, Tevfik Kosar, M. S. Q. Zulkar Nine
Large-scale graph neural network (GNN) training often requires distributed clusters because graph structure and feature tensors no longer fit in a single node's memory. In sampling…
GreenDyGNN: Runtime-Adaptive Energy-Efficient Communication for Distributed GNN Training
Arefin Niam, Tevfik Kosar, M. S. Q. Zulkar Nine
Distributed GNN training is dominated by remote feature fetching, which can be very costly. Multi-hop neighborhood sampling crosses partition boundaries and triggers fine-grained R…
A Periodic Space of Distributed Computing: Vision & Framework
Mohsen Amini Salehi, Adel N. Tousi, Hai Duc Nguyen +5
Advances in networking and computing technologies throughout the early decades of the 21st century have transformed long-standing dreams of pervasive communication and computation…
Energy-Efficient and High-Performance Data Transfers with DRL Agents
Hasibul Jamil, Jacob Goldverg, Elvis Rodrigues +2
The rapid growth of data across fields of science and industry has increased the need to improve the performance of end-to-end data transfers while using the resources more efficie…
RapidGNN: Energy and Communication-Efficient Distributed Training on Large-Scale Graph Neural Networks
Arefin Niam, Tevfik Kosar, M S Q Zulkar Nine
Graph Neural Networks (GNNs) have become popular across a diverse set of tasks in exploring structural relationships between entities. However, due to the highly connected structur…
EMLIO: Minimizing I/O Latency and Energy Consumption for Large-Scale AI Training
Hasibul Jamil, MD S Q Zulkar Nine, Tevfik Kosar
Large-scale deep learning workloads increasingly suffer from I/O bottlenecks as datasets grow beyond local storage capacities and GPU compute outpaces network and disk latencies. W…