Search papers, labs, and topics across Lattice.
Affiliation:
4
0
7
FlashQuant achieves up to 4.18x faster outlier-aware LLM inference by fusing dense and sparse computations, addressing a critical bottleneck in memory efficiency.
Transforming context ahead of time can slash time-to-first-token by nearly 12x, revolutionizing LLM agent efficiency.
Training-time data augmentation can slash validation loss in language model pretraining, making it feasible to train effectively on limited datasets for hundreds of epochs.
Forget GPU-centric designs: AMMA slashes attention latency by 15x and energy consumption by 7x with a memory-centric architecture for long-context LLMs.