From the 3 of 127 papers with an AI index.
55 citations
- University of ChicagoUS54 papers
- Brookhaven National LaboratoryUS44 papers
- SLAC National Accelerator LaboratoryUS40 papers
- Michigan State UniversityUS38 papers
- Lawrence Berkeley National LaboratoryUS36 papers
- The Ohio State UniversityUS32 papers
- Université Paris-SaclayFR30 papers
- University of California SystemUS30 papers
- Indiana UniversityUS29 papers
- University College LondonGB29 papers
- University of MichiganUS29 papers
- Czech Technical University in PragueCZ28 papers
6 papers · 1 filter
Application Failures and Machine Computational Efficiency
Carlo Graziani, Bethany Lusch, O. E. Bronson Messer
We present a framework for evaluating uptime efficiency of Exascale-class scientific computers when application failure rates are appreciable. This is the situation that confronts…
Towards Transparent Checkpointing with AI-driven Code Generation
Hai Duc Nguyen, Tekin Bicer, Kyle Chard +2
Adding reliable checkpoint/restart support to an MPI scientific application is a time-consuming expert effort that requires deep knowledge of both the application and resilience. W…
StreamGuard: Low-Overhead Resilience for Real-time HPC Data Streams
Hai Duc Nguyen, Bogdan Nicolae, Tekin Bicer +4
Real-time scientific workflows operate on continuous data streams and must produce timely, high-quality results despite executing on complex, failure-prone infrastructure. Hardware…
Implementing True MPI Sessions and Evaluating MPI Initialization Scalability
Hui Zhou, Kenneth Raffenetti, Yanfei Guo +2
Sessions is one of the major features introduced in the MPI-4 standard. It offers an alternative to the traditional world communicator model by allowing applications to construct c…
Coordinated Power Management on Heterogeneous Systems
Zhong Zheng, Zhiling Lan, Xingfu Wu +2
Performance prediction is essential for energy-efficient computing in heterogeneous computing systems that integrate CPUs and GPUs. However, traditional performance modeling method…
STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems
Chris Egersdoerfer, Philip Carns, Shane Snyder +2
I/O performance is crucial to efficiency in data-intensive scientific computing; but tuning large-scale storage systems is complex, costly, and notoriously manpower-intensive, maki…