Search papers, labs, and topics across Lattice.
Mechanistic Interpretability
2
0
3
Steering a language model's intertemporal preferences can induce significant shifts in decision-making, impacting how AI systems advise on delayed costs and benefits.
Interpolating between opposing directorial personas not only reveals surprising coherence improvements but also uncovers a shared moral-tone substrate in transformer models.