Search papers, labs, and topics across Lattice.
Affiliation:
4
0
8
2
Steering persona features can amplify emergent misalignment rates in language models beyond what traditional fine-tuning achieves.
Training LLMs with a unified moral-value dataset can enhance their alignment with human values without sacrificing general performance.
Hybrid architectures that combine attention and recurrence can maintain reasoning performance as task complexity increases, while transformers see a sharp performance drop-off.
LLMs' apparent Theory of Mind evaporates when tasks are slightly perturbed, and Chain-of-Thought prompting, surprisingly, can make things worse.