14 papers
MEMPOWER: Efficient Power Management with Fine-grained Memory Analysis and Modeling for HPC Workloads
Nanda Velugoti, Joseph Manzano, Andres Marquez +2
Managing the energy consumption and power efficiency of parallel applications is a significant issue in both HPC environments and in the cloud. As emerging applications continue to…
NOMAD: Generating Embeddings for Massive Distributed Graphs
Aishwarya Sarkar, Sayan Ghosh, Nathan R. Tallent +1
Successful machine learning on graphs or networks requires embeddings that not only represent nodes and edges as low-dimensional vectors but also preserve the graph structure. Esta…
Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training
Cunyang Wei, Siddharth Singh, Aishwarya Sarkar +7
Graph neural networks (GNNs) are widely used for learning on graph datasets derived from various real-world scenarios. Learning from extremely large graphs requires distributed tra…
SCOPE: Semantic Coreset with Orthogonal Projection Embeddings for Federated learning
Md Anwar Hossen, Nathan R. Tallent, Luanzheng Guo +1
Scientific discovery increasingly requires learning on federated datasets, fed by streams from high-resolution instruments, that have extreme class imbalance. Current ML approaches…
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
Md Hasanur Rashid, Jesun Firoz, Nathan R. Tallent +3
With the increasing importance of distributed scientific workflows, there is a critical need to ensure Quality of Service (QoS) constraints, such as minimizing time or limiting exe…
Rudder: Steering Prefetching in Distributed GNN Training using LLM Agents
Aishwarya Sarkar, Sayan Ghosh, Nathan Tallent +3
Large-scale Graph Neural Networks (GNNs) are typically trained by sampling a vertex's neighbors to a fixed distance. Because large input graphs are distributed, training requires f…