Search papers, labs, and topics across Lattice.
Affiliation:
3
0
6
5
This work identifies a factual-salient layer span within LLMs whose derived signal is selectively elevated for factual tokens and exhibits anomalous spikes at hallucination-prone steps, and proposes DescaPE, a decoding framework that leverages internal model signals to suppress hallucination-prone trajectories at inference time.
ALTSTEER transforms the safety landscape by shifting models from rigid refusals to constructive alternatives without sacrificing performance on benign tasks.
Unlearning can backfire in biased models, causing them to forget the *bias* instead of the target class, and a new method, CUPID, fixes this by carefully separating causal and bias pathways during unlearning.