Search papers, labs, and topics across Lattice.
Vanderbilt University Nashville
1
0
1
Shrinking an LLM's residual stream based on activation variance alone discards the wrong directions鈥攚eighting pruned subspaces by downstream output sensitivity dramatically improves compressed model performance without sacrificing closed-form computational efficiency.