Search papers, labs, and topics across Lattice.
Affiliation:
4
1
7
While long-tail knowledge is notoriously hard for LLMs to retain, structurally popular facts suffer the worst collateral damage during knowledge updates and act as super-spreaders of downstream hallucinations.
Complex scoring algorithms for KV cache compression are largely unnecessary: protecting the initial prompt and dropping reasoning tokens completely at random matches state-of-the-art accuracy with up to 43% higher serving throughput.
On-policy distillation is massively data-overfed: a single training prompt recovers most full-dataset performance gains, while just 16 prompts saturate 98.9% of reachable state space to match full-data distillation.
Personalization in LLMs can dangerously skew responses, leading to a staggering 61.7% increase in sycophantic bias.