Search papers, labs, and topics across Lattice.
King Abdullah University of Science and Technology
2
0
5
Achieving up to 13.9x speedup in LLM inference by leveraging cache-resident execution could redefine performance benchmarks for large models on standard CPUs.
Train your LLMs and ViTs faster: a nonparametric teaching method cuts training time by up to 21% *without* sacrificing accuracy.