Search papers, labs, and topics across Lattice.
This paper introduces the Lit2Test benchmark, designed to evaluate research proposals generated by large language models by requiring them to specify falsifiable outcomes. By organizing proposals around a six-field contract, the benchmark allows for a clear and decidable assessment of idea quality, moving beyond subjective judgments. The results reveal that the quality of proposed tests and metrics, rather than mere fluency, is critical for distinguishing between the capabilities of four leading models across 10,000 bootstrap replicates.
Language models can now be rigorously evaluated on their ability to generate falsifiable research ideas, not just stylistic fluency.
Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper rewards recovery of one realized trajectory. We introduce a benchmark that carries a proposal from Literature to Test: the Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable. Built prospectively from 200 real-paper neighborhoods, Lit2Test elicits proposals from four frontier models and compares them through 1,200 pairwise comparisons judged blind in both presentation orders. The protocol audits its own reliability through diagnostic controls and bounded human calibration, with three annotators corroborating the conclusions within explicitly stated reliability bounds. Lit2Test recovers a strict ranking of the four models in all 10,000 bootstrap replicates, and the separation comes from the quality of the proposed tests and metrics rather than from surface fluency. We release the benchmark, construction pipeline, and audit artifacts for public use.