Search papers, labs, and topics across Lattice.
This paper introduces HeteroPanacea, a simulation framework that evaluates the benefits of disaggregated serving architectures for agentic LLM inference, focusing on the distinct computational needs of prefill, decode, attention, and feedforward network (FFN) stages. The research highlights that disaggregating these components can lead to a significant increase in serving throughput, achieving up to a 75% improvement compared to traditional homogeneous GPU systems. Key findings indicate that a four-way disaggregation of the prefill, decode, attention, and FFN stages consistently enhances throughput across various model architectures, particularly when utilizing custom NPUs.
Disaggregating LLM inference stages can boost throughput by up to 75%, reshaping how we design future AI hardware systems.
Agentic inference now dominates the LLM inference landscape, requiring LLMs to actively engage in multi-turn interactions with tool-calling capabilities. This introduces a more complex workload for the underlying inference system: serving stages such as prefill and decode exhibit substantially different behaviors and demand distinct compute and memory-bandwidth capabilities. As a result, a single homogeneous GPU system now struggles to support agentic inference, motivating an industry shift toward heterogeneous systems with disaggregated serving capabilities, such as the emerging Vera-Rubin platform with GPUs and Groq LPUs. However, the question of what the optimal hardware should look like for each component in a heterogeneous system remains underexplored. To this end, we propose a novel simulation framework for disaggregated serving, termed \textbf{HeteroPanacea}, that enables system-level simulation across three dimensions: 1) disaggregated quantization, 2) automated intra- and inter-device parallelization scheduling, and 3) PDAF (prefill-decode-attention-FFN) NPU architectural heterogeneity. By combining these three axes, we provide a cross-stack simulation framework for future heterogeneous agentic serving systems. We confirm the benefit of Prefill Decode disaggregation, simulating increased serving throughput by up to 75\% compared to traditional serving with current GPUs and demonstrate 4 way Prefill Decode Attention FFN disaggregation is the most consistent for increasing throughput across different models, assuming custom NPUs. We also investigate the relationship between model architecture and gain from disaggregation by running a set of ablation studies.