Search papers, labs, and topics across Lattice.
School of Computing and Augmented Intelligence, Arizona State University
1
0
1
0
Implicit reasoning steering can covertly amplify latent biases in language models, shifting their predictions through seemingly innocuous text.