Search papers, labs, and topics across Lattice.
Affiliation:
5
0
5
3
HBF can significantly boost LLM serving efficiency by enabling more expert replicas and reducing loading times, all while preserving critical execution paths.
CODA redefines edge video diffusion by achieving up to 1.80x faster inference and 1.74x better energy efficiency without sacrificing output quality.
By intelligently managing resource contention, ADS-Tile slashes processing waste to below 1.2%, revolutionizing how DNNs are scheduled in autonomous driving systems.
Unlocking the potential of compute-in-memory accelerators for LLMs requires carefully navigating a complex dataflow design space, and AccelCIM provides the first systematic framework to do so.
A hybrid-bonding-based LLM serving accelerator, Helios, tackles the dynamic nature of KV cache management in LLM serving, achieving significant speedup and energy efficiency gains over existing GPU/NMP designs.