Search papers, labs, and topics across Lattice.
This paper explores the adaptation of the AX kernel from the Nekbone HPC mini-application to the Tenstorrent Wormhole RISC-V accelerator, focusing on the efficient evaluation of the Poisson operator. The authors identify host-side data transposition as a significant bottleneck in performance and propose two optimization strategies that lead to substantial improvements. The results demonstrate that the Wormhole accelerator achieves 242.97 GFLOPS for 100,000 elements across 128 Tensix cores, outperforming a 24-core Xeon Platinum CPU while consuming approximately seven times less power.
Achieving 242.97 GFLOPS on a RISC-V accelerator not only outpaces traditional CPUs but also does so with a fraction of the power consumption.
The growing availability of commodity RISC-V hardware has sparked interest in its use for High Performance Computing (HPC), with PCIe accelerator cards offering a practical near-term pathway to adoption. The Tenstorrent Wormhole is one example, with dedicated vector and matrix units across 128 Tensix cores, and is widely available. In this paper, we explore porting the AX kernel of Nekbone, a widely used HPC mini-application derived from the Gordon Bell Prize-winning Nek5000 spectral element solver, onto the Wormhole accelerator. This kernel evaluates the Poisson operator, and we describe the mapping of the algorithm onto the Tensix. The initial performance results reveal that the host-side data transposition, required for the z-direction gradient computation, is a severe bottleneck. Consequently, we investigated two optimisation strategies that yield dramatic improvements, achieving 242.97 GFLOPS for 100000 elements across 128 Tensix cores, outperforming a 24-core Xeon Platinum CPU and drawing approximately 7 times less power.