- Preprint Article
- 10.5194/egusphere-egu22-13484
Preparing NEMO4.2, the new NEMO modelling framework for the next generation HPCinfrastructures
- Mar 28, 2022
- Francesca Mele + 4 more +4
<p>Nowadays one of the main challenges in scientific computational field is developing the next<br>generation of HPC technologies, applications and systems towards exascale. This leads to focus the<br>efforts on the development of a new, efficient, stable and scalable NEMO reference code with<br>improved performances adapted to exploit future HPC technologies in the context of CMEMS systems.<br>On the main factors that limit the current scalability is an inefficient exploitation of the single node<br>performance. Different technical solutions have been tested to fully exploit memory hierarchies and<br>hardware peak performance. Between all, the fusion of DO loops together and by dividing the<br>computation over tiles are the two optimization strategies more efficiently take advantage of the<br>cache memory organization. This work focuses on the first one.<br>The loop fusion is a transformation which takes two adjacent loops that have the same iteration space<br>traversal and combines their bodies into a single loop. This optimization improves data locality so<br>giving a better exploitation of the cache memory and a reduction of the memory footprint because the<br>temporary arrays can be replaced with scalar values.<br>Performance tests have been executed on a domain size of 3002x2002x31 grid points running 1-year<br>GYRE_PISCES simulations with IO disabled on the Zeus Intel Xeon Gold 6154 machine, available at<br>CMCC. An increasing number of cores - from 504 to 2016 – have been used to test experiments with<br>the different HPC options.<br>The analysis focused on the routines where the optimizations have been applied. The use of the<br>extended halo introduces a penalty in the execution time that grows as the number of processes<br>increases and generally the use of loop fusion optimization slightly improves the performance. For<br>many routines, as subdomains get smaller, the improvements due to optimizations are less significant.<br>The simultaneous application of all optimizations leads to an improvement between 10% and 50%<br>(except for lateral diffusion). Looking at the total elapsed time, the new HPC optimizations speed up </p><p>the elapsed time of a factor 1.25x. Unfortunately, non-optimized routines mitigate this improvement.</p><p>The same scalability test has been repeated running 1-month ORCA025 simulations with the output<br>set to be produced at the end of the run. The results show that the use of loop fusion optimization<br>slightly improves the performance. The use of tiling in ORCA025 introduces less benefits with<br>reference to GYRE. The simultaneous application of all optimizations doesn't lead many benefits in<br>ORCA025 since the improvement concerns only a subset of routines with the Tracer lateral diffusion<br>routine getting worse in all cases.<br>In conclusion the impact of the new optimized code behaves differently depending on the<br>configuration. The overhead introduced by the extended halo implies a computation time cost that the<br>proposed optimizations are able to regain difficultly. Tiling is the aspect with the highest impact in<br>these optimizations (especially in GYRE) and loop fusion has in general a low impact. The optimizations<br>should be applied to all the rest of the code to obtain more benefits.</p>
Read more