Search papers, labs, and topics across Lattice.
This paper introduces Pallas, a proactive key-value (KV) cache migration framework designed to enhance large language model (LLM) inference during cellular handovers in AI-RAN environments. By predicting the target base station and preparing the inference state in advance, Pallas significantly reduces service interruption time (SIT) and inter-token latency (ITL) compared to traditional methods. The results show that Pallas achieves a reduction in average SIT by factors ranging from 2.28 to 89.68 and lowers ITL by 16.0% to 50.0% across multiple LLMs and varying inter-gNB link speeds.
Pallas cuts service interruption time by up to 89.68 times during LLM inference handovers, transforming mobile AI experiences.
AI-RAN brings large language model (LLM) serving close to mobile users, but cellular handover can separate an active request from its inference state: the user attaches to a target base station (gNB) while the large and growing key-value (KV) cache remains at the source. Retaining inference at the source preserves service continuity but persistently increases inter-token latency (ITL), whereas recovering the state at the target restores serving locality but requires KV-cache transfer, recomputation, or a combination of both only after handover, directly prolonging service interruption time (SIT). This work presents Pallas, a \textit{proactive} KV-cache migration framework that prepares the inference state at the predicted target before handover, in parallel with ongoing source-side inference and token delivery. At the preparation trigger, Pallas partitions the token sequence into a stable historical prefix and an evolving suffix. The target reconstructs the prefix through local prefill, while the source streams the KV blocks generated for the suffix. At handover, the target assembles both portions into an up-to-date KV cache and resumes decoding locally, leaving only unfinished preparation to contribute to SIT. An online scheduler selects the \textit{prefetching window}, which determines how early preparation begins before handover, based on mobility predictions and runtime telemetry. Across three LLMs and $100$--$500~\mathrm{Mbps}$ inter-gNB links, our vLLM-based prototype reduces average SIT by factors of $2.28$--$89.68$ over target-side recovery approaches and lowers average ITL by $16.0\%$--$50.0\%$ compared with source-side forwarding.