Search papers, labs, and topics across Lattice.
3
0
5
0
XMerge outperforms existing depth-compression techniques, achieving top rankings in task performance while enabling aggressive layer removal without sacrificing quality.
Larger AI models are more prone to catastrophic failures from stale memory, with over-reliance on outdated facts reaching near-perfect rates in certain scenarios.
Global allocation of precision budget in LLMs can improve accuracy by 21-52 points compared to local layer-specific repairs, defying conventional wisdom about quantization damage.