Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
Standard AI benchmarks often measure evaluation formatting rather than their purported constructs鈥攚ith "reasoning" and "knowledge" metrics failing to discriminate from one another, and bias benchmarks correlating more strongly with general reasoning than with peer bias tests.
Iterative problem-solving reveals that LLMs, while less accurate, can generate insightful solutions that illuminate failure modes in student coding attempts.