Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
13
WnW reduces GPU memory usage to 20% of audio tokens without sacrificing accuracy, challenging the limitations of existing KV cache methods in long-form speech processing.