Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
1
Pairwise rewards in reinforcement learning can significantly boost the robustness of LLM auditors, enhancing their ability to detect hidden model behaviors with minimal false positives.