2 papers
cs.DC2025
ML-Based Optimum Number of CUDA Streams for the GPU Implementation of the Tridiagonal Partition Method
Milena Veneva, Toshiyuki Imamura
This paper presents a heuristic for finding the optimum number of CUDA streams by using tools common to the modern AI-oriented approaches and applied to the parallel partition algo…
cs.MS2024
Interface for Sparse Linear Algebra Operations
Ahmad Abdelfattah, Willow Ahrens, Hartwig Anzt +32
The standardization of an interface for dense linear algebra operations in the BLAS standard has enabled interoperability between different linear algebra libraries, thereby boosti…