Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
Prefill latency can be slashed by 80% without sacrificing long-context retrieval accuracy by pairing concatenation-aware fine-tuning with selective KV recomputation.
Mitigating the bias introduced by Top-K retrieval in long-context transformers can be achieved by estimating the contribution of unretrieved tokens with a fixed-size feature map, significantly improving performance without increasing KV-cache reads.