Search papers, labs, and topics across Lattice.
The Hong Kong Polytechnic University
1
0
3
Unlock practical memory savings and faster decoding in LLMs with PrunePath, a structured sparsification method that adaptively activates experts based on token-level probability budgets.