Search papers, labs, and topics across Lattice.
2
1
6
2
Traditional load balancing in EP MoE serving can lead to inefficiencies, but a new makespan-aware dispatcher achieves up to 15.5% throughput gains by adapting to varying compute and memory constraints.
Memory optimization that prioritizes belief clarity over mere outcome success can radically enhance long-horizon reasoning in LLMs.