Search papers, labs, and topics across Lattice.
Department of Mechanical Engineering, Santa Clara University, Santa Clara, CA 95053, USA
1
0
3
0
Fusing gather, GEMM, and scatter operations in a single CUDA kernel slashes SIMP topology optimization runtime by up to 7.3x, proving that optimized kernel fusion can still yield massive gains even in memory-bound scientific computing.