Search papers, labs, and topics across Lattice.
Zhejiang University, Hangzhou High-Tech Zone (Binjiang, Institute of Blockchain and Data Security
1
0
2
LLM embedding models waste over 40% of their inference compute processing prefix tokens that become completely redundant at deeper layers, enabling aggressive, training-free token pruning with virtually zero retrieval degradation.