most citedHiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights

8 citations · 9 across the 9 of their papers we have counts for

collaborators

12 papers

cs.DC2025

FIRST: Federated Inference Resource Scheduling Toolkit for Scientific AI Model Access

Aditya Tanikanti, Benoit Côté, Yanfei Guo +9

We present the Federated Inference Resource Scheduling Toolkit (FIRST), a framework enabling Inference-as-a-Service across distributed High-Performance Computing (HPC) clusters. FI…

cs.LG20251 cited

AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions

Väinö Hatanpää, Eugene Ku, Jason Stock +12

Generative machine learning offers new opportunities to better understand complex Earth system dynamics. Recent diffusion-based methods address spectral biases and improve ensemble…

cs.LG2025

PagedEviction: Structured Block-wise KV Cache Pruning for Efficient Large Language Model Inference

Krishna Teja Chitty-Venkata, Jie Ye, Xian-He Sun +4

KV caching significantly improves the efficiency of Large Language Model (LLM) inference by storing attention states from previously processed tokens, enabling faster generation of…

cs.LG2025

LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference

Krishna Teja Chitty-Venkata, Sandeep Madireddy, Murali Emani +1

Mixture-of-Experts (MoE) models scale efficiently by activating only a subset of experts per token, offering a computationally sparse alternative to dense architectures. While prio…

cs.DC2025

Aurora: Architecting Argonne's First Exascale Supercomputer for Accelerated Scientific Discovery

William E. Allcock, Benjamin S. Allen, James Anchell +106

Aurora is Argonne National Laboratory's pioneering Exascale supercomputer, designed to accelerate scientific discovery with cutting-edge architectural innovations. Key new technolo…

cs.LG2025

MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models

Krishna Teja Chitty-Venkata, Sylvia Howland, Golara Azar +5

Mixture of Experts (MoE) models have enabled the scaling of Large Language Models (LLMs) and Vision Language Models (VLMs) by achieving massive parameter counts while maintaining c…