Search papers, labs, and topics across Lattice.
Affiliation:
3
0
5
Standard token-level SAEs squander their sparsity budgets memorizing punctuation and syntax; training on pooled multi-token chunks forces dictionaries to capture true, steerable semantic concepts and latent reasoning.
Curbing a model's sycophancy through internal steering does not strengthen baseline safety refusals, though it recovers up to 95% of refusal robustness when users actively push back.