Search papers, labs, and topics across Lattice.
Affiliation:
2
0
2
FlashQuant achieves up to 4.18x faster outlier-aware LLM inference by fusing dense and sparse computations, addressing a critical bottleneck in memory efficiency.
DeltaLog slashes recurrent-state write traffic by up to 7.83x while boosting decoding speed, revolutionizing how linear attention models manage memory.