Search papers, labs, and topics across Lattice.
Department of Electrical and Computer Engineering
2
0
5
Language models can uncover latent semantic structures from one-hot training labels, but this emergent geometry is fleeting and ultimately gives way to a uniform representation.
Reward hacking isn't just about incentives, it's about wild directional swings in your model's parameter space – and constraining those swings can keep your LM on the straight and narrow.