From the 1 of 383 papers with an AI index.
2.6k citations
- A. Ishikawa2 profiles64 · h 72
- V. Gaur2 profiles63 · h 51
- V. Zhilich2 profiles63 · h 68
- H. Hayashii2 profiles62 · h 75
- J. Libby2 profiles62 · h 58
- P. Krokovny2 profiles62 · h 0
- S. Nishida2 profiles62 · h 54
- D. Liventsev2 profiles61 · h 71
- E. Solovieva2 profiles61 · h 53
- H. Aihara2 profiles61 · h 85
- T. Sumiyoshi2 profiles61 · h 76
- E. Won2 profiles60 · h 74
- Virginia TechUS82 papers
- University of Hawaii SystemUS74 papers
- Chinese Academy of SciencesCN71 papers
- University of CincinnatiUS71 papers
- High Energy Accelerator Research OrganizationJP70 papers
- Karlsruhe Institute of TechnologyDE70 papers
- Charles UniversityCZ69 papers
- Korea Institute of Science & Technology InformationKR69 papers
- The University of TokyoJP69 papers
- Technical University of MunichDE68 papers
- Indian Institute of Technology MadrasIN67 papers
- University of Hawaiʻi at MānoaUS67 papers
18 papers · 1 filter
LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models
Minqiu Sun, Xin Huang, Luanzheng Guo +3
Checkpointing is essential for fault tolerance in training large language models (LLMs). However, existing methods, regardless of their I/O strategies, periodically store the entir…
Scrutinizing Variables for Checkpoint Using Automatic Differentiation
Xin Huang, Weiping Zhang, Shiman Meng +4
Checkpoint/Restart (C/R) saves the running state of the programs periodically, which consumes considerable system resources. We observe that not every piece of data is involved in…
Distributed Order Recording Techniques for Efficient Record-and-Replay of Multi-threaded Programs
Xiang Fu, Shiman Meng, Weiping Zhang +6
After all these years and all these other shared memory programming frameworks, OpenMP is still the most popular one. However, its greater levels of non-deterministic execution mak…
Concepts for designing modern C++ interfaces for MPI
C. Nicole Avans, Alfredo A. Correa, Sayan Ghosh +5
Since the C++ bindings were deleted in 2008, the Message Passing Interface (MPI) community has revived efforts in building high-level modern C++ interfaces. Such interfaces are eit…
MAPA: Multi-Accelerator Pattern Allocation Policy for Multi-Tenant GPU Servers
Kiran Ranganath, Joshua D. Suetterlein, Joseph B. Manzano +2
Multi-accelerator servers are increasingly being deployed in shared multi-tenant environments (such as in cloud data centers) in order to meet the demands of large-scale compute-in…
Efficient Exascale Discretizations: High-Order Finite Element Methods
Tzanio Kolev, Paul Fischer, Misun Min +27
Efficient exploitation of exascale architectures requires rethinking of the numerical algorithms used in many large-scale applications. These architectures favor algorithms that ex…