Search papers, labs, and topics across Lattice.
Cohere Labs
2
0
4
Confidence calibration in reasoning models can be dramatically improved by aligning supervision with the model's state, cutting ECE by over half on challenging benchmarks.
Decomposing complex tasks into verifiable checklists unlocks more effective reinforcement learning, but only if you can avoid the pitfalls of reward hacking and verifier bias.