Search papers, labs, and topics across Lattice.
Beihang University
2
0
3
Accepting selective mismatches in autoregressive decoding can boost throughput by over 15% without any additional training or model adjustments.
ReTopK achieves a 3.07x speedup in attention computation with only a 0.50% increase in perplexity, revolutionizing long-context processing efficiency.