Resource Oblivious Sorting on Multicores
arXiv:1508.01504 · doi:10.1145/3040221
Abstract
We present a deterministic sorting algorithm, SPMS (Sample, Partition, and Merge Sort), that interleaves the partitioning of a sample sort with merging. Sequentially, it sorts elements in time cache-obliviously with an optimal number of cache misses. The parallel complexity (or critical path length) of the algorithm is , which improves on previous bounds for optimal cache oblivious sorting. The algorithm also has low false sharing costs. When scheduled by a work-stealing scheduler in a multicore computing environment with a global shared memory and cores, each having a cache of size organized in blocks of size , the costs of the additional cache misses and false sharing misses due to this parallel execution are bounded by the cost of and cache misses respectively, where is the number of steals performed during the execution. Finally, SPMS is resource oblivious in Athat the dependence on machine parameters appear only in the analysis of its performance, and not within the algorithm itself.
A version very similar to this appears in ACM Transactions on Parallel Computing (TOPC), Vol. 3, No. 4, Article 23, 2017. The current version adds some additional citations to earlier sorting algorithms, and a comparison to Sharesort
References in corpus (1)
Cited by in corpus (7)
- The Parallel Persistent Memory Model
- Many Sequential Iterative Algorithms Can Be Parallel and (Nearly) Work-efficient
- Low-Depth Parallel Algorithms for the Binary-Forking Model without Atomics
- Efficient Stepping Algorithms and Implementations for Parallel Shortest Paths
- Bounding Cache Miss Costs of Multithreaded Computations Under General Schedulers
- Analysis of Work-Stealing and Parallel Cache Complexity
- Balanced Partitioning of Several Cache-Oblivious Algorithms