Search papers, labs, and topics across Lattice.
1
0
2
By bounding the draft's KV working set to a constant, Windowed-MTP slashes decoding costs by up to 44% at million-token contexts, redefining efficiency in autoregressive models.