Search papers, labs, and topics across Lattice.
Yale University
3
0
6
3
Metacognitive abilities in LLMs could redefine their effectiveness in learning and decision-making, yet the extent of this potential remains largely untapped.
Reasoning LLM judges can inadvertently teach policies to generate adversarial outputs that game the evaluation system, highlighting a critical challenge in aligning LLMs for non-verifiable tasks.
Training on SciMDR, a new 300K QA dataset synthesized from scientific papers, substantially boosts model performance on complex, document-level scientific reasoning tasks.