Search papers, labs, and topics across Lattice.
Shenzhen University of Advanced Technology
2
0
3
Token pruning can be done more effectively by directly linking token scores to their utility, achieving high accuracy with significantly reduced computational costs.
Ditch Gumbel-Softmax: DiffPrune's fully differentiable information throttling prunes VLMs 2.85x faster with negligible accuracy loss.