Search papers, labs, and topics across Lattice.
3
0
5
0
Achieving up to 2.35× faster inference with only 19–28% of the original KV cache, VisCache redefines efficiency in Vision Large Language Models.
Achieving superior inference accuracy and resource efficiency in multi-cluster edge intelligence networks hinges on a delicate balance between model pruning and collaborative inference.
Fine-tune LLMs on your phone without sacrificing privacy or blowing up your data plan: pAirZero slashes memory and communication costs by using zeroth-order optimization over the air.