3 citations · 3 across the 3 of their papers we have counts for
4 papers
A Doubly-pipelined, Dual-root Reduction-to-all Algorithm and Implementation
Jesper Larsson Träff
We discuss a simple, binary tree-based algorithm for the collective allreduce (reduction-to-all, MPI_Allreduce) operation for parallel systems consisting of suitably interconne…
-ported vs. -lane Broadcast, Scatter, and Alltoall Algorithms
Jesper Larsson Träff
In -ported message-passing systems, a processor can simultaneously receive different messages from other processors, and send different messages to other process…
Stamp-it: A more Thread-efficient, Concurrent Memory Reclamation Scheme in the C++ Memory Model
Manuel Pöter, Jesper Larsson Träff
We present Stamp-it, a new, concurrent, lock-less memory reclamation scheme with amortized, constant-time (thread-count independent) reclamation overhead. Stamp-it has been impleme…
A Note on (Parallel) Depth- and Breadth-First Search by Arc Elimination
Jesper Larsson Träff
This note recapitulates an algorithmic observation for ordered Depth-First Search (DFS) in directed graphs that immediately leads to a parallel algorithm with linear speed-up for a…