Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
TailSieve achieves up to 2.59x speedup in LLM rollouts by intelligently routing long-tail requests, transforming how we handle high-concurrency decoding.