Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
Predictive delta prefetching can achieve up to 12x faster LLM inference on edge devices, even when models exceed memory limits.