Par Do IT Now

Next-Gen Velocity. Benchmarking the Amazon EC2 Hpc8a Evolution.

Introduction

Do IT Now deployed an HPC cluster in the cloud using AMD-based HPC-optimized instances on AWS, in order to benchmark the next generation of Amazon EC2 Hpc instances.

This study compares Amazon EC2 Hpc8a preview instances with Amazon EC2 Hpc7a instances across representative open-source HPC workloads. The objective was to evaluate performance, scalability, and configuration behavior across different node and core counts, while identifying the tuning parameters required to make the most of the underlying architecture.

During this benchmark campaign, Do IT Now also explored optimization patterns related to MPI configuration, process pinning, memory bandwidth usage, and multi-node scaling in a cloud-based HPC environment.

Objective

The objective of this project is to provide benchmark results for Amazon EC2 Hpc8a preview instances and compare them with the previous-generation Amazon EC2 Hpc7a instances.

Both instance types are based on AMD EPYC™ processors and are designed for demanding HPC workloads. The comparison focuses on two major open-source applications commonly used in engineering and scientific computing:

  • OpenRadioss™, used for explicit finite element analysis workloads.
  • OpenFOAM™, used for computational fluid dynamics workloads.

Do IT Now used two comparable instance configurations to ensure a fair and accurate comparison:

  • Amazon EC2 Hpc8a preview instances, featuring 5th Gen AMD EPYC™ processors, 192 physical cores per node, 768 GiB RAM, DDR5 memory, and 300 Gbps networking.
  • Amazon EC2 Hpc7a instances, featuring 4th Gen AMD EPYC™ processors, 192 physical cores per node, 768 GiB RAM, DDR5 memory, and 300 Gbps networking.

The goal was not only to measure raw performance, but also to understand how each application behaves depending on core count, memory bandwidth, interconnect usage, and multi-node scaling.

About the software

The software used for the benchmarks is open source:

OpenRadioss™ is an analysis solution used to evaluate and optimize product performance for highly nonlinear problems under dynamic loadings. It is widely used across industrial sectors for use cases such as crashworthiness, safety, and manufacturability of complex designs.

OpenFOAM™ is a general-purpose Computational Fluid Dynamics solver capable of addressing demanding industrial and scientific applications. Its scalable solver technology makes it a relevant candidate for evaluating cloud-based HPC performance across different node configurations.

About the models

All the models used in this study are well suited to benchmarking HPC performance on clusters with a large number of CPUs.

OpenRadioss™

For OpenRadioss™, Do IT Now used the publicly available Taurus 10 million finite elements model, recovered from the OpenRadioss™ public repository.

The model includes three simulation settings based on simulation time:

  • Full simulation: 120 milliseconds.
  • Shorter simulation: 10 milliseconds, best suited for performance and scalability testing.
  • Very short simulation: 2 milliseconds, useful for testing HPC cluster functionality.

For this benchmark, Do IT Now used the 10-millisecond simulation test case.

OpenFOAM™

For OpenFOAM™, Do IT Now used polyMesh models with different mesh resolutions to compare execution times under different configurations.

The benchmark includes:

  • A Coarse model with 65 million cells.
  • A Fine model with 236 million cells.

 

These test cases were used to analyze how OpenFOAM behaves across single-node and multi-node configurations, and to evaluate the impact of memory bandwidth and core density on solver performance.

Conclusions

The benchmark campaign shows that optimal HPC configuration is highly solver-dependent.

For OpenRadioss™, the Hpc8a instances deliver a clear performance advantage over Hpc7a across all tested configurations. The benchmark demonstrates strong scaling up to 384 cores, which appears to be the optimal configuration for this specific workload. Beyond that point, scalability becomes less efficient, mainly due to communication overhead and workload-specific decomposition limits.

For OpenFOAM™, both Hpc7a and Hpc8a demonstrate consistent scalability as the cluster expands. However, Hpc8a shows a significant performance improvement compared to Hpc7a, with runtime reductions of roughly one third across the tested scenarios.

The results also highlight an important architectural consideration: high core density does not automatically lead to the best performance for every workload. While OpenRadioss™ benefits from the available core count, OpenFOAM™ can experience memory bandwidth saturation when fully populating a node. In this case, distributing fewer cores across more nodes can improve performance by increasing the available aggregate memory bandwidth.

Overall, Amazon EC2 Hpc8a instances show a strong performance uplift compared to Hpc7a and confirm their relevance for large-scale MPI workloads. The study also reinforces the importance of careful HPC tuning, including MPI settings, process pinning, memory management, and workload-specific configuration choices.

Discuss your HPC needs today

Contact us at info@doitnowgroup.com

Partager

Lire d'autres documents techniques