Acceleration of multi-component multiple-precision arithmetic with branch-free algorithms and SIMD vectorization
arXiv:2603.14926 · doi:10.1007/978-3-032-30521-3_32
Abstract
Multiple-precision floating-point branch-free algorithms can significantly accelerate multi-component arithmetic implemented by combining hardware-based binary64 and binary32, particularly for triple- and quadruple-precision computations. In this study, we achieved benchmark results on x86 and ARM CPU platforms to quantify the accelerations achieved in linear computations and polynomial evaluation by integrating these algorithms.