Iterative and Parallel Performance Analysis of Non-Blocking Communication Algorithms in the Massively Parallel Neutron Transport Code PIDOTS
- 1. North Carolina State University Department of Nuclear Engineering 2146 Burlington Engineering Laboratory Raleigh, NC, USA, 27603 (United States)
- 2. Los Alamos National Laboratory (United States)
Description
The PIDOTS neutral particle transport code utilizes a red/black implementation of the Parallel Gauss-Seidel algorithm to solve the SN approximation of the neutron transport equation on 3D Cartesian meshes. PIDOTS is designed for execution on massively parallel platforms and is capable of using the full resources of modern, leadership class high performance computers. Initial testing revealed that some configurations of PIDOTS's Integral Transport Matrix Method solver demonstrated unexpectedly poor parallel scaling. Work at Idaho and Los Alamos National Laboratories then revealed that this inefficiency was a result of the accumulation of high-cost latency events in the complex blocking communication networks employed during each PIDOTS iteration. That work explored the possibility of minimizing those inefficiencies while maintaining a blocking communications model. While significant speedups were obtained, it was shown that fully mitigating the problem on general-purpose platforms was highly unlikely for a blocking code. This work continues that analysis by implementing a deeply interleaved non-blocking communication model into PIDOTS. This new model benefits from the optimization work performed on the blocking model while also providing significant opportunities to overlap the remaining un-mitigated communication costs with computation. Additionally, our new approach is easily transferable to other similarly spatially decomposed codes. The resulting algorithm was tested on LANL's Trinity system at up to 32,768 processors and was found at that processor count to effectively hide 100% of MPI communication cost – equivalently 20% of the red/black phase time. It is expected that the implemented interleaving algorithm can fully support far higher processor counts and completely hide communication costs up ~50% of total iteration time.
Availability note (English)
Available from https://www.epj-conferences.org/articles/epjconf/pdf/2021/01/epjconf_physor2020_03016.pdf; https://doaj.org/article/668ddb38c3bc4e02a49d9ed181000700Additional details
Identifiers
Publishing Information
- Journal Title
- EPJ. Web of Conferences
- Journal Volume
- 247
- Journal Page Range
- vp.
- ISSN
- 2100-014X
Conference
- Title
- International Conference on Physics of Reactors: Transition to a Scalable Nuclear Future
- Acronym
- PHYSOR2020
- Dates
- 28 Mar - 2 Apr 2020
- Place
- Cambridge (United Kingdom)
INIS
- Country of Publication
- France
- Country of Input or Organization
- France
- INIS RN
- 53087984
- Subject category
- S22: GENERAL STUDIES OF NUCLEAR REACTORS; S97: MATHEMATICAL METHODS AND COMPUTING;
- Resource subtype / Literary indicator
- Conference
- Descriptors DEI
- ALGORITHMS; DESIGN; ITERATIVE METHODS; MATRICES; NEUTRON TRANSPORT; NEUTRON TRANSPORT THEORY; OPTIMIZATION; PERFORMANCE; TESTING
- Descriptors DEC
- CALCULATION METHODS; MATHEMATICAL LOGIC; NEUTRAL-PARTICLE TRANSPORT; RADIATION TRANSPORT; TRANSPORT THEORY