Search papers, labs, and topics across Lattice.
Email:
1
0
2
Quantizing small language models can be predictable and efficient, with a novel metric that identifies optimal layers for speed and quality trade-offs.