Search papers, labs, and topics across Lattice.
Huazhong University of Science and Technology 搂 Xiaohongshu Inc
2
0
5
FlowBlock achieves up to 4.01脳 faster decoding and 77.1% lower latency in dLLMs by transforming block dependencies into scheduling resources through innovative wavefront-parallel techniques.
LLM external memory performance degrades as the memory grows across rounds in realistic streaming scenarios, highlighting the need for better memory management strategies.