Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
Achieving 2.27x faster decoding for 2-bit LLM weights could redefine efficiency benchmarks in low-bit quantization methods.