Search papers, labs, and topics across Lattice.
Affiliation:
1
0
1
Heuristic KV cache pruning routinely fails under heavy compression because it ignores softmax curvature, but treating attention as a nonlinear Gaussian channel reveals exactly which tokens preserve information capacity.