Search papers, labs, and topics across Lattice.
Affiliation:
1
0
3
FlashAttention-V achieves up to 42x speedup in transformer inference on CPUs, transforming how we leverage vector architectures for small language models.