Search papers, labs, and topics across Lattice.
Huazhong University of Science and Technology 搂 Xiaohongshu Inc
2
0
4
FlowBlock achieves up to 4.01脳 faster decoding and 77.1% lower latency in dLLMs by transforming block dependencies into scheduling resources through innovative wavefront-parallel techniques.
Save 20% on LLM costs with <2% accuracy drop by strategically cascading a small model with a large one, guided by a confidence-calibrated SLM.