Search papers, labs, and topics across Lattice.
University of Pennsylvania
2
0
5
ConflictScore reveals that language models often overlook conflicting evidence, leading to overconfident and inaccurate claims.
RLVR's reasoning gains hinge on high-entropy tokens, revealing a critical inefficiency in uniform reward broadcast that EAPO effectively addresses.