Search papers, labs, and topics across Lattice.
Affiliation:
2
0
4
TailSieve achieves up to 2.59x speedup in LLM rollouts by intelligently routing long-tail requests, transforming how we handle high-concurrency decoding.
TAILS resolves cross-task ambiguities in continual learning by directly correcting feature representations, leading to significant performance boosts without changing the underlying model.