Search papers, labs, and topics across Lattice.
2
0
3
Standard CoT prompting may distract advanced LLMs from their reasoning tasks, leading to worse performance than zero-shot approaches.
Transforming uninformative reward signals into actionable insights, Reasoning Arena boosts reasoning performance while slashing training costs.