Search papers, labs, and topics across Lattice.
This paper addresses the "generalization crisis" in matrix multiplication (MatMul) for dynamic tensor shapes on Ascend NPUs by introducing AdaptCore, an adaptive framework that optimizes performance through spatial tiling and instruction orchestration. By mapping dynamic shapes into a hardware-aware 2D tiling taxonomy, AdaptCore effectively balances on-chip capacity and multi-core parallelism while leveraging a composable optimization library and a deterministic analytical performance model. The results indicate a significant 1.85x mean speedup across 80,000 input shapes, outperforming existing vendor libraries by up to 1.48x in end-to-end model acceleration.
Achieving a 1.85x speedup in matrix multiplication on Ascend NPUs could redefine performance benchmarks for dynamic tensor operations.
Matrix Multiplication (MatMul) faces a"generalization crisis"driven by highly dynamic tensor shapes. This crisis is particularly acute on Ascend NPUs, where explicitly controlled architectures and strict physical constraints render existing GPU-centric optimizations ineffective. To resolve this, we propose AdaptCore, an adaptive framework for universally high-performance MatMul on Ascend NPUs. AdaptCore systematically decouples operator optimization into spatial tiling and instruction orchestration. It first maps dynamic shapes into a hardware-aware 2D tiling taxonomy to balance on-chip capacity limits and multi-core parallelism. Furthermore, it integrates a composable optimization library with a deterministic analytical performance model. By mathematically evaluating hardware state mutations, AdaptCore proactively selects and caches optimal implementations, enabling O(1) overhead runtime dispatching. Evaluations demonstrate that AdaptCore delivers a remarkable 1.85x mean speedup across 80,000 input shapes, and achieves up to a 1.48x acceleration in representative end-to-end models over the highly-tuned native vendor library (ACLNN).