paper

A performance enhancement of the Payne-Hanek range reduction algorithm

arXiv:2609.35015

Abstract

Range reduction plays a crucial role in the accuracy and performance of evaluating trigonometric functions, and is often the primary bottleneck for large floating-point inputs. While fast algorithms such as Cody--Waite work efficiently over narrow intervals, the Payne--Hanek algorithm remains the standard technique for accurate reduction across large floating-point inputs. However, many existing implementations of Payne--Hanek suffer from high latency due to heavy branching, conversion overheads, and the use of multi-word integer arithmetic, which hinders SIMD vectorization. In this paper, we analyze and present a branch-free variation of the Payne--Hanek algorithm using only floating-point arithmetic. Our method operates directly over large double-precision inputs () and is well suited to hardware with FMA instructions. We formulate the precision constraints in terms of a truncation error budget, construct a compact lookup table indexed by the input exponent, and prove that the scaled reduced argument has absolute error below and relative error below for every input, including worst cases. The same routine can serve both as the complete range reduction of a single-stage implementation and as the fast path of a correctly rounded one, achieving higher throughput than existing implementations and lower latency than those returning a double-double reduced argument. The algorithm is currently implemented in the LLVM libc project.

Updated to include CORE-MATH newer Payne-Hanek range reduction implementation, and add CRLIBM to performance analysis. Updated the performance table after fixing the setup