Search papers, labs, and topics across Lattice.
2
0
3
11
ExaGEMM achieves a staggering 13.29x latency reduction for low-bit GEMM on CPUs, revolutionizing how we approach efficient ML inference.
PolyQ achieves up to 32.1% better perplexity at a 3-bit target while reducing activation reorder traffic by nearly 71%, proving that fractional-bit quantization can be both practical and efficient for CPU inference.