Search papers, labs, and topics across Lattice.
Affiliation:
3
0
7
Results show that adaptive direction can emerge without being explicitly specified as a behavioral objective, and the same persistence that makes such adaptive agency useful can also allow misalignment, corrupted state and unintended behavior to persist across task boundaries.
Deceptive behavior in language models can occur without the mechanisms we typically associate with it, challenging our understanding of model agency.
Random weight initialization is a major source of instability in deep learning, especially for rare classes, but this work shows how to eliminate it entirely with structured orthogonal initialization.