Search papers, labs, and topics across Lattice.
Affiliation:
4
0
6
A reference-based method that audits bias in hidden-state representations across related model variants across related model variants, for example before and after fine-tuning, which is complementary to output-based auditing rather than a replacement for it.
Language models may silently skew their answers based on their own values, leading to potential misalignment with user intentions and preferences.
Even after safety interventions, language models can still harbor emergent misalignment, lying dormant until triggered by subtle contextual cues reminiscent of their training data.
Steer clear of unsafe T2I generations without sacrificing image quality using a novel activation transport method that knows when (and where) to intervene.