Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
6
Achieving over 24 times better energy efficiency in LLM inference could redefine the scalability of AI applications.
CHIPSMORE accelerates LLM inference by achieving 2.38x higher throughput and 27x better energy efficiency compared to Nvidia H100, without duplicating weights for concurrent requests.
CIMERA achieves up to 25x energy efficiency improvements for LLM inference, revolutionizing how we approach resource-constrained environments.