Search papers, labs, and topics across Lattice.
Nanjing University
2
0
4
High variance in draft quality is tackled head-on, leading to a remarkable 66% increase in throughput for speculative decoding in large language models.
JetSpec shatters the speed ceiling of speculative decoding, achieving up to 9.64x acceleration on complex tasks while maintaining high acceptance rates.