Search papers, labs, and topics across Lattice.
Affiliation:
2
0
4
Most privacy-enhancing technologies merely secure the pipelines for inherently harmful functionalities rather than preventing the harms themselves.
Weak-to-strong reward models can ace the test but still fail in the real world, revealing a hidden brittleness in current preference learning approaches.