Search papers, labs, and topics across Lattice.
Affiliation:
2
0
5
0
Pairwise rewards in reinforcement learning can significantly boost the robustness of LLM auditors, enhancing their ability to detect hidden model behaviors with minimal false positives.
Alignment auditing tools that shine in isolation can completely flop when deployed in an autonomous agent, revealing a critical gap in current evaluation methodologies.