Search papers, labs, and topics across Lattice.
1
0
3
This paper proposes OmniKVQuant, a training-free framework that enables 2-bit KV caches while substantially preserving performance across seven audio-visual benchmarks and provides a fused Triton decode kernel that unpacks the 2-bit cache during attention, so no dense FP16 cache is ever built.