论文信息 - Impacts of Multi-GPU MPI Collective Communications on Large FFT Computation

Impacts of Multi-GPU MPI Collective Communications on Large FFT Computation

Most applications targeting exascale, such as those part of the Exascale Computing Project (ECP), are designed for heterogeneous architectures and rely on the Message Passing Interface (MPI) as their underlying parallel programming model. In this paper we analyze the limitations of collective MPI communication for the computation of fast Fourier transforms (FFTs), which are relied on heavily for large-scale particle simulations. We present experiments made at one of the largest heterogeneous platforms, the Summit supercomputer at ORNL. We discuss communication models from state-of-the-art FFT libraries, and propose a new FFT library, named HEFFTE (Highly Efficient FFTs for Exascale), which supports heterogeneous architectures and yields considerable speedups compared with CPU libraries, while maintaining good weak as well as strong scalability.

[1] Jack Dongarra,et al. GPUDirect MPI Communications and Optimizations to Accelerate FFTs on Exascale Systems , 2019 .

[2] Jack J. Dongarra,et al. Towards dense linear algebra for hybrid GPU accelerated manycore systems , 2009, Parallel Comput..

[3] Hal Finkel,et al. HACC , 2016, Commun. ACM.

[4] Mei Han An,et al. accuracy and stability of numerical algorithms , 1991 .

[5] James Demmel,et al. Communication-avoiding algorithms for linear algebra and beyond , 2013, 2013 IEEE 27th International Symposium on Parallel and Distributed Processing.

[6] Jack Dongarra,et al. Evaluation and Design of FFT for Distributed Accelerated Systems , 2018 .

[7] Steven G. Johnson,et al. The Design and Implementation of FFTW3 , 2005, Proceedings of the IEEE.

[8] Jack Dongarra,et al. Design and Implementation for FFT-ECP on Distributed Accelerated Systems , 2019 .

[9] J. Dongarra,et al. ECP Milestone Report FFT-ECP Implementation Optimizations and Features Phase WBS 2 . 3 . 3 . 09 , Milestone FFT-ECP ST-MS-10-1440 Stanimire , 2019 .