Search papers, labs, and topics across Lattice.
2
0
4
3
Thinking tokens may create an illusion of deliberation in reasoning models, but they often lock in decisions early, undermining safety efforts.
Turn sparse binary rewards into dense supervision signals by having a model revise its own work, then distilling the revision strategy back into the original generation.