Search papers, labs, and topics across Lattice.
This paper introduces the Marvell Photonic Fabric Memory Appliance, a novel photonic-CXL hybrid architecture designed to address the memory limitations in large language model (LLM) inference. By replacing traditional electrical switches with a passive fiber shuffle, the system achieves 32 TB of shared memory across 16 hosts, resulting in over 50% latency reduction compared to existing electrical CXL pools. The findings reveal a significant improvement in time-to-first-token by 6.6x for multi-turn conversation workloads, effectively eliminating cache eviction issues and enabling more efficient memory management for concurrent long-context users.
A photonic memory architecture can slash LLM inference latency by over 50%, paving the way for scalable, high-performance KV cache management.
LLM inference at scale faces a memory wall. The KV cache demands tens of terabytes at hundreds of gigabytes per second, yet no current memory tier delivers both at once. Characterization across multi-generation GPU systems with various LLaMA models shows host memory retrieval achieves up to 100x speedup over re-computation but supports only tens of concurrent long-context users. Electrical CXL pooling theoretically bridges this gap, but switch latency, cable reach limits, and power-scaling issues prevent practical TB-scale deployments. We present the Marvell Photonic Fabric Memory Appliance, a photonic-CXL hybrid architecture replacing electrical switches with a passive fiber shuffle to deliver 32 TB shared memory across 16 hosts via a switch-free full- crossbar topology. Emulation results demonstrate over 50 percent latency reduction versus electrical CXL pools. Simulation results show that the PF Memory Appliance eliminates cache eviction cliffs by improving time-to-first-token by 6.6x for multi-turn conversations workloads.