Search papers, labs, and topics across Lattice.
Baidu Inc
2
0
5
Learned attention allocation patterns reveal that SWA is best positioned in lower layers, challenging conventional wisdom on attention distribution in LLMs.
Forget brute-force hinting: KnowRL distills knowledge into atomic units, then uses subset selection to find the *least* amount of guidance needed to supercharge LLM reasoning.