Search papers, labs, and topics across Lattice.
Affiliation:
7
0
8
0
APT accelerates high-resolution diffusion models by up to 8.16脳 while enhancing energy efficiency, revolutionizing the feasibility of real-time generative AI applications.
Forget-Retain Alignment Gap reveals that the structure of weight updates, not just their distance, is key to preventing LLMs from relearning forgotten information.
NOVA's innovative architecture delivers 4.5x higher throughput and 69.8% lower latency for hybrid LLMs, pushing the boundaries of memory processing efficiency.
Physical isolation in chiplet design can boost LLM serving performance by nearly 50% while slashing latency by over 60%.
LLMs can run over 2x faster on your phone without any extra training, thanks to a clever pruning and quantization strategy that doesn't sacrifice accuracy.
Diffusion models can run 4.7x faster and use 68% less energy thanks to a new sparsity-aware hardware accelerator that reuses cached tokens and attention sparsity patterns.
Video diffusion models can be accelerated by 4.5x with a novel hardware-software co-design that leverages output activations to guide token reduction, achieving significantly higher compression ratios than prior art.