Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
2
This work presents a tensor-based formulation of the Viterbi algorithm for HSMMs, restructuring the inner loops into tensor operations that naturally map onto SIMD units and massively parallel architectures and provides optimized implementations spanning single- and multi-core CPUs, and, for the first time, GPU.
Understanding how concurrent AI training jobs disrupt each other could redefine resource allocation strategies in supercomputing environments.
HPC interconnects choke under AI-like bursty traffic, but the degree of performance degradation varies wildly across different fabrics like InfiniBand, Slingshot, and Ethernet.