Search papers, labs, and topics across Lattice.
Affiliation:
1
0
1
Fixed-size chunked prefill forces an unnecessary compromise between decode latency and launch overhead: dynamically sizing chunks to fit active decode deadlines boosts serving goodput by up to 3.3脳 under tight latency SLOs.