3 papers
cs.PF2023
Cache Optimization and Performance Modeling of Batched, Small, and Rectangular Matrix Multiplication on Intel, AMD, and Fujitsu Processors
Sameer Deshmukh, Rio Yokota, George Bosilca
Factorization and multiplication of dense matrices and tensors are critical, yet extremely expensive pieces of the scientific toolbox. Careful use of low rank approximation can dra…
math.NA2023
distributed direct factorization of structured dense matrices using runtime systems
Sameer Deshmukh, Qinxiang Ma, Rio Yokota +1
Structured dense matrices result from boundary integral problems in electrostatics and geostatistics, and also Schur complements in sparse preconditioners such as multi-frontal met…
math.NA2022
Scalable Linear Time Dense Direct Solver for 3-D Problems Without Trailing Sub-Matrix Dependencies
Qianxiang Ma, Sameer Deshmukh, Rio Yokota
Factorization of large dense matrices are ubiquitous in engineering and data science applications, e.g. preconditioners for iterative boundary integral solvers, frontal matrices in…