Search papers, labs, and topics across Lattice.
1
0
2
Achieving a 2.00x reduction in weight bandwidth while decoding at 5.94 tokens/s could redefine efficiency benchmarks for autoregressive models on CPU architectures.