Search papers, labs, and topics across Lattice.
Peking University
3
0
6
12
ExpertPlex slashes over 95% of duplicate model weights and boosts goodput by up to 2.01 times, transforming how we serve large language models.
Agentic RL rollouts are bottlenecked by long-tail trajectory generation, but Heddle's trajectory-centric approach achieves 2.5x higher throughput.
Double your LLM inference throughput by routing KV-cache through decoding engines to bypass the bandwidth bottleneck on prefill engines.