collaborators

14 papers

cs.DC2026

MEMPOWER: Efficient Power Management with Fine-grained Memory Analysis and Modeling for HPC Workloads

Nanda Velugoti, Joseph Manzano, Andres Marquez +2

Managing the energy consumption and power efficiency of parallel applications is a significant issue in both HPC environments and in the cloud. As emerging applications continue to…

cs.LG2026

NOMAD: Generating Embeddings for Massive Distributed Graphs

Aishwarya Sarkar, Sayan Ghosh, Nathan R. Tallent +1

Successful machine learning on graphs or networks requires embeddings that not only represent nodes and edges as low-dimensional vectors but also preserve the graph structure. Esta…

cs.LG2026

Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training

Cunyang Wei, Siddharth Singh, Aishwarya Sarkar +7

Graph neural networks (GNNs) are widely used for learning on graph datasets derived from various real-world scenarios. Learning from extremely large graphs requires distributed tra…

cs.LG2026

SCOPE: Semantic Coreset with Orthogonal Projection Embeddings for Federated learning

Md Anwar Hossen, Nathan R. Tallent, Luanzheng Guo +1

Scientific discovery increasingly requires learning on federated datasets, fed by streams from high-resolution instruments, that have extreme class imbalance. Current ML approaches…

cs.DC2026

QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models

Md Hasanur Rashid, Jesun Firoz, Nathan R. Tallent +3

With the increasing importance of distributed scientific workflows, there is a critical need to ensure Quality of Service (QoS) constraints, such as minimizing time or limiting exe…

cs.LG2026

Rudder: Steering Prefetching in Distributed GNN Training using LLM Agents

Aishwarya Sarkar, Sayan Ghosh, Nathan Tallent +3

Large-scale Graph Neural Networks (GNNs) are typically trained by sampling a vertex's neighbors to a fixed distance. Because large input graphs are distributed, training requires f…