Search papers, labs, and topics across Lattice.
4
0
4
AngelSpec achieves nearly double the inference speed of traditional methods while improving output quality by intelligently adapting to the specific demands of different tasks.
D-Cut transforms speculative decoding efficiency by cutting verification costs, achieving up to 3.0x speedup over traditional methods in high-concurrency scenarios.
DFlare achieves up to 5.52x speedup in LLM inference by allowing draft layers to independently leverage richer target knowledge, breaking through previous capacity constraints.
Speculative decoding gets a throughput boost of up to 4.32x by using reinforcement learning to dynamically balance drafting and verification.