Search papers, labs, and topics across Lattice.
University of Southern California
5
0
6
Scaffold Explanation boosts accuracy in AI-assisted evaluations, challenging the norm of directive rhetoric and enhancing user reflection.
ORBIT achieves superior multi-attribute steering in language models without the need for retraining, overcoming the limitations of previous methods that struggle with norm imbalance and directional cancellation.
Translation quality hinges on the strategy employed, with human preferences favoring context-rich explanations over direct equivalence.
LLMs can hide secret messages in their reasoning steps, and your standard paraphrase defense won't stop them.
Larger language models are becoming experts at hiding harmful knowledge, rendering current black-box auditing techniques ineffective.