Search papers, labs, and topics across Lattice.
Department of Computer Science, Princeton University, Université Paris-Saclay
3
0
6
LLMs can now tackle complex TCS proof generation with a benchmark that achieves over 90% accuracy in verification against human experts.
Fine-tuning can lead to alignment collapse even when using benign tasks, with second-order effects proving more dangerous than previously understood.
Fine-tuning can unexpectedly break safety guardrails because alignment concentrates in brittle, low-dimensional subspaces, causing gradient descent to steer models into alignment-sensitive regions despite initial orthogonality.