Search papers, labs, and topics across Lattice.
1
0
2
Reusing KV caches during model swaps can retain up to 98% of prefill accuracy while running 2.7-25x faster than traditional methods.