Exact diagonalization of quantum lattice models on coprocessors
arXiv:1511.00863 · doi:10.1016/j.cpc.2016.07.018
Abstract
We implement the Lanczos algorithm on an Intel Xeon Phi coprocessor and compare its performance to a multi-core Intel Xeon CPU and an NVIDIA graphics processor. The Xeon and the Xeon Phi are parallelized with OpenMP and the graphics processor is programmed with CUDA. The performance is evaluated by measuring the execution time of a single step in the Lanczos algorithm. We study two quantum lattice models with different particle numbers, and conclude that for small systems, the multi-core CPU is the fastest platform, while for large systems, the graphics processor is the clear winner, reaching speedups of up to 7.6 compared to the CPU. The Xeon Phi outperforms the CPU with sufficiently large particle number, reaching a speedup of 2.5.
References in corpus (9)
- Nearly-flat bands with nontrivial topology
- Fractional quantum Hall effect in the absence of Landau levels
- Lattice QCD with Domain Decomposition on Intel Xeon Phi Co-Processors
- Phase Transition in 3d Heisenberg Spin Glasses with Strong Random Anisotropies, through a Multi-GPU Parallelization
- Computational Physics on Graphics Processing Units
- Exact diagonalization of the Hubbard model on graphics processing units
- Quantum phases of disordered flatband lattice fractional quantum Hall systems
- Ageing at the Spin-Glass/Ferromagnet Transition: Monte Carlo Simulation using GPUs
- Impurities and Landau level mixing in a fractional quantum Hall state in a flatband lattice model