Search papers, labs, and topics across Lattice.
This paper introduces TailSieve, a framework designed to optimize the routing and allocation of long-tail generation requests in large language model (LLM) rollouts. By utilizing partial rollouts to identify candidate tail groups and employing a hierarchical controller for dynamic routing adjustments, TailSieve achieves significant improvements in makespan and throughput. The method demonstrates up to 2.59x speedup over traditional uniform routing approaches, highlighting its effectiveness in managing long-tail generation requests in high-concurrency environments.
TailSieve achieves up to 2.59x speedup in LLM rollouts by intelligently routing long-tail requests, transforming how we handle high-concurrency decoding.
Large-scale rollouts have become a core component of modern LLM systems, spanning reinforcement learning (RL) post-training, on-policy distillation (OPD), and sampling-heavy evaluation pipelines. Unlike online serving, which is typically optimized for request-level latency and throughput, a small number of long-tail generations can dominate the end-to-end makespan of an entire rollout step. In practice, rollout requests are often routed uniformly across replicas, which can place extremely long generations inside high-concurrency decoding batches. To address this, we present TailSieve, a partial-rollout-guided framework that jointly controls tail routing and replica allocation for LLM rollouts. In an idealized setting with known completion lengths, we show that makespan-optimal routing in the long-tail regime combines tail isolation with load balancing, and that a simple top-k policy closely approximates this offline optimum. Leveraging the observation that long-tail prompts tend to remain long-tailed across policy updates, TailSieve uses partial rollouts as a training-free signal for identifying candidate tail groups. A hierarchical controller then jointly adapts the number of isolated groups and the replica split between the tail and bulk pools using collected response-work history and a measured concurrency-throughput model. TailSieve achieves up to 1.67x routing-only speedup over uniform group routing. The resulting low-concurrency tail pool further enables route-specialized speculative decoding with MTP or DFlash, achieving up to 2.59x speedup over uniform routing. Selected prompts are regenerated under the current policy, preserving on-policy generation and avoiding additional routing-induced length bias in steady state.