Search papers, labs, and topics across Lattice.
Affiliation:
2
0
4
3
FlashAttention-V achieves up to 42x speedup in transformer inference on CPUs, transforming how we leverage vector architectures for small language models.
Energy-efficient multi-GPU scheduling can be significantly enhanced by jointly optimizing task placement and DVFS, yielding up to 14.8% energy savings without sacrificing performance.