Search papers, labs, and topics across Lattice.
This study systematically characterizes the scalability and performance of large-scale AI training on modern high-performance computing (HPC) systems, focusing on configurations with up to 2400 GPUs. By evaluating various parallelization strategies and their interaction with network congestion and compute capabilities, the authors quantify communication overheads and analyze how concurrent training jobs interfere with one another through a realistic noise model. The findings reveal critical insights into execution efficiency and scalability, providing a benchmark suite that enhances our understanding of AI workload performance in multi-tenant environments.
Understanding how concurrent AI training jobs disrupt each other could redefine resource allocation strategies in supercomputing environments.
Characterising AI workload performance on modern HPC systems requires understanding both their scalability in isolation and their behaviour under concurrent execution. However, the interplay among parallelisation strategies, network congestion, compute capability, and interconnect technologies remains poorly understood. This work investigates the performance and scalability of AI models up to 2400 GPUs. We quantify the communication overheads and their impact across different interconnects by evaluating scale-up, scale-out, and rack-scale configurations under multiple allocation schemes. Finally, we study how multiple concurrent training jobs interfere with each other by designing a realistic noise model. We design a benchmark suite of AI models to evaluate the performance of five distinct parallelisation strategies across different supercomputing clusters, including Alps, Leonardo, LUMI, JUPITER, NVL72 GB300, and DGX A100. Our work provides a systematic characterization of the scalability and execution efficiency of distributed AI training, while offering key insights into performance behavior under realistic multi-tenant scenarios.