2 citations · 7 across the 17 of their papers we have counts for
22 papers
ExaServe: Large-Scale LLM Serving on Exascale HPC Systems
Wenyi Wang, Shu Shi, Yadu Babuji +2
Cloud-native LLM serving frameworks have made deployment routine in data centers, yet deploying them on leadership-class supercomputers remains an engineering challenge requiring s…
Avatar: Toward Autonomous End-to-End Orchestration of Scientific Workflows using LLMs
Suman Raj, Hai Duc Nguyen, Haochen Pan +3
Scientific workflow management (WMSs) systems automate execution, yet orchestrate using fixed, hand-tuned rules. LLM agents promise more autonomous orchestration, but it remains un…
Diamond Agent: Agentic Control of Federated HPC Resources as a Service
Haotian Xie, Junlin Chen, Mingkai Zheng +9
Efficiently aggregating and orchestrating computing power across heterogeneous clusters for HPC workflows faces four practical challenges: preserving workflow context across indepe…
StreamGuard: Low-Overhead Resilience for Real-time HPC Data Streams
Hai Duc Nguyen, Bogdan Nicolae, Tekin Bicer +4
Real-time scientific workflows operate on continuous data streams and must produce timely, high-quality results despite executing on complex, failure-prone infrastructure. Hardware…
When More Cores Hurts: The Vector Database Scaling Paradox in HPC
Seth Ockerman, Song Young Oh, Amal Gueroudji +12
Vector databases have been designed and optimized for cloud environments; however, emerging scientific AI workloads (e.g., molecular search, meteorological trajectory detection, an…
Icicle: Scalable Metadata Indexing and Real-Time Monitoring for HPC File Systems
Haochen Pan, Ryan Chard, Song Young Oh +7
Modern HPC file systems can contain billions of files and hundreds of petabytes of data, making even simple questions increasingly intractable to answer. Traditional file system ut…